Fetching the paper…
Reading the bibliography…
We present the KL3M tokenizers, a family of specialized tokenizers for legal, financial, and governmental text.
J. J. Webster and C. Kit, “Tokenization as the initial phase in nlp,” in COLING 1992 volume 4: The 14th international conference on computational linguistics , 1992
1992
Earlier work this paper cites.
P. Gage, “A new algorithm for data compression,” The C Users Journal , vol. 12, no. 2, pp. 23–38, 1994
1994
Earlier work this paper cites.
R. Sennrich, B. Haddow, and A. Birch, “Neural machine translation of rare words with subword units,” in Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2016, pp. 1715–1725
2016
Earlier work this paper cites.
T. Kudo and J. Richardson, “Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations , 2018, pp. 66–71
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers) , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
I. Beltagy, K. Lo, and A. Cohan, “Scibert: A pretrained language model for scientific text,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , 2019, pp. 3615–3620
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Cited alongside, same era.
K. Bostrom and G. Durrett, “Byte pair encoding is suboptimal for language model pretraining,” in Findings of the Association for Computational Linguistics: EMNLP 2020 , 2020, pp. 4617–4624
2020
Cited alongside, same era.
J. Lee, W. Yoon, S. Kim, D. Kim, S. Kim, C. H. So, and J. Kang, “Biobert: a pre-trained biomedical language representation model for biomedical text mining,” Bioinformatics , vol. 36, no. 4, pp. 1234–1240, 2020
2020
Cited alongside, same era.
D. Q. Nguyen, T. Vu, and A. Tuan Nguyen, “Bertweet: A pre-trained language model for english tweets,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations , 2020, pp. 9–14
2020
P. Rust, J. Pfeiffer, I. Vulić, S. Ruder, and I. Gurevych, “How good is your tokenizer? on the monolingual performance of multilingual language models,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , 2021, pp. 3118–3135
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. Mansar, J. Kang, and I. E. Maarouf, “The finsim-2 2021 shared task: Learning semantic similarities for the financial domain,” in Companion Proceedings of the Web Conference 2021 , 2021, pp. 288–292
2021
Later among the works it cites.
J. H. Clark, D. Garrette, I. Turc, and J. Wieting, “Canine: Pre-training an efficient tokenization-free encoder for language representation,” Transactions of the Association for Computational Linguistics , vol. 10, pp. 73–91, 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
I. Chalkidis, M. Fergadiotis, P. Malakasiotis, N. Aletras, and I. Androutsopoulos, “Legal-bert: The muppets straight out of law school,” in Findings of the Association for Computational Linguistics: EMNLP 2020 , 2020, pp. 2898–2904
2020
Cited alongside, same era.
W. Ma, Y. Cui, C. Si, T. Liu, S. Wang, and G. Hu, “Charbert: Character-aware pre-trained language model,” in Proceedings of the 28th International Conference on Computational Linguistics , 2020, pp. 39–50
2020
Cited alongside, same era.
2022
Later among the works it cites.
C. Wang, X. Liu, Z. Chen, H. Hong, J. Tang, and D. Song, “Deepstruct: Pretraining of language models for structure prediction,” in Findings of the Association for Computational Linguistics: ACL 2022 , 2022, pp. 803–823
2022
Later among the works it cites.
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann et al. , “Palm: Scaling language modeling with pathways,” Journal of Machine Learning Research , vol. 24, no. 240, pp. 1–113, 2023
2023
Later among the works it cites.