Fetching the paper…
Reading the bibliography…
Current language models (LMs) use a fixed, static subword tokenizer.
Adv-bert: Bert is not robust on misspellings! generating nature adversarial samples on bert
Lichao Sun, Kazuma Hashimoto, Wenpeng Yin, Akari Asai, Jia Li, Philip Yu, and Caiming Xiong. 2020 · 2003
Earlier work this paper cites.
Japanese and korean voice search
Mike Schuster and Kaisuke Nakajima. 2012 · 2012
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Mimicking word embeddings using subword RNNs
Yuval Pinter, Robert Guthrie, and Jacob Eisenstein. 2017 · 2017
Earlier work this paper cites.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Earlier work this paper cites.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo. 2018 · 2018
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Earlier work this paper cites.
Breaking the softmax bottleneck: A high-rank RNN language model
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W. Cohen. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Attentive mimicking: Better word embeddings by attending to informative contexts
Timo Schick and Hinrich Schütze. 2019 · 2019
Earlier work this paper cites.
Byte pair encoding is suboptimal for language model pretraining
Kaj Bostrom and Greg Durrett. 2020 · 2020
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Earlier work this paper cites.
CharacterBERT: Reconciling ELMo and BERT for word-level open-vocabulary representations from characters
Hicham El Boukkouri, Olivier Ferret, Thomas Lavergne, Hiroshi Noji, Pierre Zweigenbaum, and Jun’ichi Tsujii. 2020 · 2020
Earlier work this paper cites.
Accelerating large-scale inference with anisotropic vector quantization
Ruiqi Guo, Philip Sun, Erik Lindgren, Quan Geng, David Simcha, Felix Chern, and Sanjiv Kumar. 2020 · 2020
Earlier work this paper cites.
BERTRAM: Improved word embeddings have big impact on contextualized model performance
Timo Schick and Hinrich Schütze. 2020 · 2020
Earlier work this paper cites.
Between words and characters: A brief history of open-vocabulary modeling and tokenization in NLP
Sabrina J Mielke, Zaid Alyafeai, Elizabeth Salesky, Colin Raffel, Manan Dey, Matthias Gallé, Arun Raja, Chenglei Si, Wilson Y Lee, Benoît Sagot, et al. 2021 · 2021
Earlier work this paper cites.
How good is your tokenizer? on the monolingual performance of multilingual language models
Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, and Iryna Gurevych. 2021 · 2021
Earlier work this paper cites.
Multi-view subword regularization
Xinyi Wang, Sebastian Ruder, and Graham Neubig. 2021 · 2021
Cited alongside, same era.
Canine: Pre-training an efficient tokenization-free encoder for language representation
Jonathan H. Clark, Dan Garrette, Iulia Turc, and John Wieting. 2022 · 2022
Cited alongside, same era.
Fast vocabulary transfer for language model compression
Leonidas Gee, Andrea Zugarini, Leonardo Rigutini, and Paolo Torroni. 2022 · 2022
Cited alongside, same era.
An embarrassingly simple method to mitigate undesirable properties of pretrained language model tokenizers
Valentin Hofmann, Hinrich Schuetze, and Janet Pierrehumbert. 2022 · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Cited alongside, same era.
XLM-V: Overcoming the vocabulary bottleneck in multilingual masked language models
Davis Liang, Hila Gonen, Yuning Mao, Rui Hou, Naman Goyal, Marjan Ghazvininejad, Luke Zettlemoyer, and Madian Khabsa. 2023 · 2023
Later among the works it cites.
Efficient transformers with dynamic token pooling
Piotr Nawrot, Jan Chorowski, Adrian Lancucki, and Edoardo Maria Ponti. 2023 · 2023
Later among the works it cites.
Impact of tokenization on language models: An analysis for turkish
Cagri Toraman, Eyup Halit Yilmaz, Furkan Şahinuç, and Oguzhan Ozcelik. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Later among the works it cites.
Why do nearest neighbor language models work?
Frank F. Xu, Uri Alon, and Graham Neubig. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Benjamin Minixhofer, Fabian Paischer, and Navid Rekabsaz. 2022 · 2022
Cited alongside, same era.
Charformer: Fast character transformers via gradient-based subword tokenization
Yi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Prakash Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu, and Donald Metzler. 2022 · 2022
Cited alongside, same era.
ByT5: Towards a token-free future with pre-trained byte-to-byte models
Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel. 2022 · 2022
Cited alongside, same era.
Do all languages cost the same? tokenization in the era of commercial language models
Orevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai, David Mortensen, Noah Smith, and Yulia Tsvetkov. 2023 · 2023
Cited alongside, same era.
Can large language models be an alternative to human evaluations?
Cheng-Han Chiang and Hung-yi Lee. 2023 · 2023
Cited alongside, same era.
FOCUS: Effective embedding initialization for monolingual specialization of multilingual models
Konstantin Dobler and Gerard de Melo. 2023 · 2023
Cited alongside, same era.
How do different tokenizers perform on downstream tasks in scriptio continua languages?: A case study in Japanese
Takuro Fujii, Koki Shibata, Atsuki Yamaguchi, Terufumi Morishita, and Yasuhiro Sogawa. 2023 · 2023
Cited alongside, same era.
MEGABYTE: predicting million-byte sequences with multiscale transformers
Lili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, and Mike Lewis. 2023 · 2023
Later among the works it cites.
Tokenizer choice for LLM training: Negligible or crucial?
Mehdi Ali, Michael Fromm, Klaudia Thellmann, Richard Rutmann, Max Lübbering, Johannes Leveling, Katrin Klug, Jan Ebert, Niclas Doll, Jasper Buschhoff, Charvi Jain, Alexander Weber, Lena Jurkschat, Hammam Abdelwahab, Chelsea John, Pedro Ortiz Suarez, Malte Ostendorff, Samuel Weinbach, Rafet Sifa, Stefan Kesselheim, and Nicolas Flores-Herr. 2024 · 2024
Closest in time.
Getting the most out of your tokenizer for pre-training and domain adaptation
Gautier Dagan, Gabriel Synnaeve, and Baptiste Rozière. 2024 · 2024
Closest in time.
Nearest neighbor speculative decoding for llm generation and attribution
Minghan Li, Xilun Chen, Ari Holtzman, Beidi Chen, Jimmy Lin, Wen-tau Yih, and Xi Victoria Lin. 2024 · 2024
Closest in time.
MYTE: Morphology-driven byte encoding for better and fairer multilingual language modeling
Tomasz Limisiewicz, Terra Blevins, Hila Gonen, Orevaoghene Ahia, and Luke Zettlemoyer. 2024 · 2024
Closest in time.
OFA: A framework of initializing unseen subword embeddings for efficient large-scale multilingual continued pretraining
Yihong Liu, Peiqin Lin, Mingyang Wang, and Hinrich Schuetze. 2024 · 2024
Closest in time.
Universal NER: A gold-standard multilingual named entity recognition benchmark
Stephen Mayhew, Terra Blevins, Shuheng Liu, Marek Suppa, Hila Gonen, Joseph Marvin Imperial, Börje Karlsson, Peiqin Lin, Nikola Ljubešić, Lester James Miranda, Barbara Plank, Arij Riabi, and Yuval Pinter. 2024 · 2024
Closest in time.
Large Language Models: A Survey
Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao. 2024 · 2024
Closest in time.
Benjamin Minixhofer, Edoardo Maria Ponti, and Ivan Vulić. 2024 · 2024
Closest in time.
Global mmlu: Understanding and addressing cultural and linguistic biases in multilingual evaluation
Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David I Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, et al. 2024 · 2024
Closest in time.
Greed is all you need: An evaluation of tokenizer inference methods
Omri Uzan, Craig W Schmidt, Chris Tanner, and Yuval Pinter. 2024 · 2024
Closest in time.