Fetching the paper…
Reading the bibliography…
Tokenization is a critical part of modern NLP pipelines.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., and Brew, J · 1910
Earlier work this paper cites.
Improving language model of human genome for dna-protein binding prediction based on task-specific pre-training
Luo, H., Shan, W., Chen, C., Ding, P., and Luo, L · 1913
Earlier work this paper cites.
Morphological word segmentation on agglutinative languages for neural machine translation
Pan, Y., Li, X., Yang, Y., and Dong, R · 2001
Earlier work this paper cites.
Arabert: Transformer-based model for arabic language understanding
Antoun, W., Baly, F., and Hajj, H. M · 2003
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Marcus, M. P., Santorini, B., and Marcinkiewicz, M. A · 2004
Earlier work this paper cites.
Unsupervised morphology induction using morfessor
Creutz, M., Lagus, K., and Virpioja, S · 2005
Earlier work this paper cites.
Morfessor 2.0: Toolkit for statistical morphological segmentation
Smit, P., Virpioja, S., Grönroos, S.-A., and Kurimo, M · 2006
Earlier work this paper cites.
V-measure: A conditional entropy-based external cluster evaluation measure
Rosenberg, A. and Hirschberg, J · 2007
Earlier work this paper cites.
Morpho challenge 2005-2010: Evaluations and results
Kurimo, M., Virpioja, S., Turunen, V., and Lagus, K · 2010
Earlier work this paper cites.
Japanese and korean voice search
Schuster, M. and Nakajima, K · 2012
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G., and Dean, J · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C. D · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2015
Earlier work this paper cites.
Enriching word vectors with subword information
Bojanowski, P., Grave, E., Joulin, A., and Mikolov, T · 2016
Earlier work this paper cites.
A joint model of orthography and morphological segmentation
Cotterell, R., Vieira, T., and Schütze, H · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, N. Q., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R · 2016
Cited alongside, same era.
Super-convergence: Very fast training of residual networks using large learning rates
Smith, L. N. and Topin, N · 2017
Cited alongside, same era.
An evaluation of two vocabulary reduction methods for neural machine translation
Ataman, D. and Federico, M · 2018
Cited alongside, same era.
Meaningless yet meaningful: Morphology grounded subword-level NMT
Banerjee, T. and Bhattacharyya, P · 2018
Cited alongside, same era.
How much does tokenization affect neural machine translation?
Domingo, M., García-Martínez, M., Helle, A., Casacuberta, F., and Herranz, M · 2018
Cited alongside, same era.
Morfessor EM+Prune: Improved subword segmentation with expectation maximization and pruning
Grönroos, S.-A., Virpioja, S., and Kurimo, M · 2020
Later among the works it cites.
DagoBERT: Generating derivational morphology with a pretrained language model
Hofmann, V., Pierrehumbert, J., and Schütze, H · 2020
Later among the works it cites.
Evaluating various tokenizers for arabic text classification
Alyafeai, Z., AlShaibani, M. S., Ghaleb, M., and Ahmad, I · 2021
Later among the works it cites.
MorphyNet: a large multilingual database of derivational and inflectional morphology
Batsuren, K., Bella, G., and Giunchiglia, F · 2021
Later among the works it cites.
Hofmann, V., Pierrehumbert, J. B., and Schütze, H · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Subword regularization: Improving neural network translation models with multiple subword candidates
Kudo, T · 2018
Cited alongside, same era.
Morphological and language-agnostic word segmentation for NMT
Machácek, D., Vidra, J., and Bojar, O · 2018
Cited alongside, same era.
Using morphological knowledge in open-vocabulary neural language models
Matthews, A., Neubig, G., and Dyer, C · 2018
Cited alongside, same era.
Deep contextualized word representations
Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Cited alongside, same era.
Morphological zero-shot neural machine translation
Zhou, G · 2018
Cited alongside, same era.
PyTorch Lightning, 3 2019
Falcon, W. and The PyTorch Lightning team · 2019
Cited alongside, same era.
DNABERT: pre-trained Bidirectional Encoder Representations from Transformers model for DNA-language in genome
Ji, Y., Zhou, Z., Liu, H., and Davuluri, R. V · 2021
Later among the works it cites.
The effectiveness of morphology-aware segmentation in low-resource neural machine translation
Saleva, J. and Lignos, C · 2021
Later among the works it cites.
PromptSource: An integrated development environment and repository for natural language prompts
Bach, S., Sanh, V., Yong, Z. X., Webson, A., Raffel, C., Nayak, N. V., Sharma, A., Kim, T., Bari, M. S., Fevry, T., Alyafeai, Z., Dey, M., Santilli, A., Sun, Z., Ben-david, S., Xu, C., Chhablani, G., Wang, H., Fries, J., Al-shaibani, M., Sharma, S., Thakker, U., Almubarak, K., Tang, X., Radev, D., Jiang, M. T.-j., and Rush, A · 2022
Later among the works it cites.
The SIGMORPHON 2022 shared task on morpheme segmentation
Batsuren, K., Bella, G., Arora, A., Martinovic, V., Gorman, K., Žabokrtský, Z., Ganbold, A., Dohnalová, Š., Ševčíková, M., Pelegrinová, K., Giunchiglia, F., Cotterell, R., and Vylomova, E · 2022
Later among the works it cites.
An embarrassingly simple method to mitigate undesirable properties of pretrained language model tokenizers
Hofmann, V., Schütze, H., and Pierrehumbert, J · 2022
Later among the works it cites.
AlephBERT: Language model pre-training and evaluation from sub-word to sentence level
Seker, A., Bandel, E., Bareket, D., Brusilovsky, I., Greenfeld, R., and Tsarfaty, R · 2022
Later among the works it cites.
Meal: Stable and active learning for few-shot prompting, 2023
Köksal, A., Schick, T., and Schütze, H · 2023
Closest in time.
MTEB: Massive text embedding benchmark
Muennighoff, N., Tazi, N., Magne, L., and Reimers, N · 2023
Closest in time.
Impact of tokenization on language models: An analysis for turkish
Toraman, C., Yilmaz, E. H., Sahinu, F., and Ozcelik, O · 2023
Closest in time.
Are all languages equally hard to language-model?
Cotterell, R., Mielke, S. J., Eisner, J., and Roark, B · 2085
Closest in time.