Fetching the paper…
Reading the bibliography…
The popular subword tokenizers of current language models, such as Byte-Pair Encoding (BPE), are known not to respect morpheme boundaries, which affects the downstream performance of the models.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
ALBERT: A lite BERT for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 1909
Earlier work this paper cites.
Introduction to wordnet: An on-line lexical database
George A Miller, Richard Beckwith, Christiane Fellbaum, Derek Gross, and Katherine J Miller. 1990 · 1990
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
A language-independent feature schema for inflectional morphology
John Sylak-Glassman, Christo Kirov, David Yarowsky, and Roger Que. 2015 · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
A joint model of orthography and morphological segmentation
Ryan Cotterell, Tim Vieira, and Hinrich Schütze. 2016 · 2016
Earlier work this paper cites.
Neural morphological analysis: Encoding-decoding canonical segments
Katharina Kann, Ryan Cotterell, and Hinrich Schütze. 2016 · 2016
Earlier work this paper cites.
Very-large scale parsing and normalization of wiktionary morphological paradigms
Christo Kirov, John Sylak-Glassman, Roger Que, and David Yarowsky. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Unimorph 2.0: Universal morphology
Christo Kirov, Ryan Cotterell, John Sylak-Glassman, Géraldine Walther, Ekaterina Vylomova, Patrick Xia, Manaal Faruqui, Sabrina J Mielke, Arya D McCarthy, Sandra Kübler, et al. 2018 · 2018
Earlier work this paper cites.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo. 2018 · 2018
Earlier work this paper cites.
Compositional Morphology Through Deep Learning
Ekaterina Vylomova. 2018 · 2018
Earlier work this paper cites.
Wic: the word-in-context dataset for evaluating context-sensitive meaning representations
Mohammad Taher Pilehvar and Jose Camacho-Collados. 2019 · 2019
Earlier work this paper cites.
On the importance of subword information for morphological tasks in truly low-resource languages
Yi Zhu, Benjamin Heinzerling, Ivan Vulić, Michael Strube, Roi Reichart, and Anna Korhonen. 2019 · 2019
Earlier work this paper cites.
Byte pair encoding is suboptimal for language model pretraining
Kaj Bostrom and Greg Durrett. 2020 · 2020
Earlier work this paper cites.
Finding the optimal vocabulary size for neural machine translation
Thamme Gowda and Jonathan May. 2020 · 2020
Earlier work this paper cites.
The morph as a minimal linguistic form
Martin Haspelmath. 2020 · 2020
Earlier work this paper cites.
Dynamic programming encoding for subword segmentation in neural machine translation
Xuanli He, Gholamreza Haffari, and Mohammad Norouzi. 2020 · 2020
Cited alongside, same era.
Compositionality decomposed: how do neural networks generalise?
Dieuwke Hupkes, Verna Dankers, Mathijs Mul, and Elia Bruni. 2020 · 2020
Cited alongside, same era.
An empirical study of tokenization strategies for various korean nlp tasks
Kyubyong Park, Joohong Lee, Seongbo Jang, and Dawoon Jung. 2020 · 2020
Cited alongside, same era.
Mind your inflections! improving nlp for non-standard englishes with base-inflection encoding
Samson Tan, Shafiq Joty, Lav Varshney, and Min-Yen Kan. 2020 · 2020
Cited alongside, same era.
Morphynet: a large multilingual database of derivational and inflectional morphology
Khuyagbaatar Batsuren, Gábor Bella, and Fausto Giunchiglia. 2021 · 2021
Cited alongside, same era.
Analyzing cognitive plausibility of subword tokenization
Lisa Beinborn and Yuval Pinter. 2023 · 2023
Later among the works it cites.
A taxonomy and review of generalization research in nlp
Dieuwke Hupkes, Mario Giulianelli, Verna Dankers, Mikel Artetxe, Yanai Elazar, Tiago Pimentel, Christos Christodoulopoulos, Karim Lasri, Naomi Saphra, Arabella Sinclair, Dennis Ulmer, Florian Schottmann, Khuyagbaatar Batsuren, Kaiser Sun, Koustuv Sinha, Leila Khalatbari, Maria Ryskina, Rita Frieske, Ryan Cotterell, and Zhijing Jin. 2023 · 2023
Later among the works it cites.
Morphpiece: Moving away from statistical language representation
Haris Jabbar. 2023 · 2023
Later among the works it cites.
Compoundpiece: Evaluating and improving decompounding performance of language models
Benjamin Minixhofer, Jonas Pfeiffer, and Ivan Vulić. 2023 · 2023
Later among the works it cites.
Words, subwords, and morphemes: What really matters in the surprisal-reading time relationship?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2021 · 2021
Cited alongside, same era.
Superbizarre is not superb: Derivational morphology improves bert’s interpretation of complex words
Valentin Hofmann, Janet Pierrehumbert, and Hinrich Schütze. 2021 · 2021
Cited alongside, same era.
Between words and characters: a brief history of open-vocabulary modeling and tokenization in nlp
Sabrina J Mielke, Zaid Alyafeai, Elizabeth Salesky, Colin Raffel, Manan Dey, Matthias Gallé, Arun Raja, Chenglei Si, Wilson Y Lee, Benoît Sagot, et al. 2021 · 2021
Cited alongside, same era.
The SIGMORPHON 2022 shared task on morpheme segmentation
Khuyagbaatar Batsuren, Gábor Bella, Aryaman Arora, Viktor Martinovic, Kyle Gorman, Zdeněk Žabokrtský, Amarsanaa Ganbold, Šárka Dohnalová, Magda Ševčíková, Kateřina Pelegrinová, Fausto Giunchiglia, Ryan Cotterell, and Ekaterina Vylomova. 2022a · 2022
Cited alongside, same era.
SIGMORPHON 2022 shared task on morpheme segmentation submission description: Sequence labelling for word-level morpheme segmentation
Leander Girrbach. 2022 · 2022
Cited alongside, same era.
Improving tokenisation by alternative treatment of spaces
Edward Gow-Smith, Harish Tayyar Madabushi, Carolina Scarton, and Aline Villavicencio. 2022 · 2022
Cited alongside, same era.
An embarrassingly simple method to mitigate undesirable properties of pretrained language model tokenizers
Valentin Hofmann, Hinrich Schütze, and Janet Pierrehumbert. 2022 · 2022
Cited alongside, same era.
Sathvik Nair and Philip Resnik. 2023 · 2023
Later among the works it cites.
Tokenization consistency matters for generative models on extractive nlp tasks
Kaiser Sun, Peng Qi, Yuhao Zhang, Lan Liu, William Wang, and Zhiheng Huang. 2023 · 2023
Later among the works it cites.
Alpaca: A strong, replicable instruction-following model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023 · 2023
Later among the works it cites.
Incorporating context into subword vocabularies
Shaked Yehezkel and Yuval Pinter. 2023 · 2023
Later among the works it cites.
Tokenization and the noiseless channel
Vilém Zouhar, Clara Meister, Juan Gastaldi, Li Du, Mrinmaya Sachan, and Ryan Cotterell. 2023 · 2023
Later among the works it cites.
Using contextual information for sentence-level morpheme segmentation
Prabin Bhandari and Abhishek Paudel. 2024 · 2024
Closest in time.
On catastrophic inheritance of large foundation models
Hao Chen, Bhiksha Raj, Xing Xie, and Jindong Wang. 2024 · 2024
Closest in time.
Unpacking tokenization: Evaluating text compression and its correlation with model performance
Omer Goldman, Avi Caciularu, Matan Eyal, Kris Cao, Idan Szpektor, and Reut Tsarfaty. 2024 · 2024
Closest in time.
Tokenization is more than compression
Craig W Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine, Omri Uzan, Yuval Pinter, and Chris Tanner. 2024 · 2024
Closest in time.
Tokenization counts: the impact of tokenization on arithmetic in frontier llms
Aaditya K Singh and DJ Strouse. 2024 · 2024
Closest in time.
Aarohi Srivastava and David Chiang. 2024 · 2024
Closest in time.
Revisiting subword tokenization: A case study on affixal negation in large language models
Thinh Hung Truong, Yulia Otmakhova, Karin Verspoor, Trevor Cohn, and Timothy Baldwin. 2024 · 2024
Closest in time.
Greed is all you need: An evaluation of tokenizer inference methods
Omri Uzan, Craig W Schmidt, Chris Tanner, and Yuval Pinter. 2024 · 2024
Closest in time.