Fetching the paper…
Reading the bibliography…
Language models can largely benefit from efficient tokenization.
The meaning-frequency relationship of words
George Kingsley Zipf. 1945 · 1945
Earlier work this paper cites.
A new algorithm for data compression
Philip Gage. 1994 · 1994
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Statistical Significance Tests for Machine Translation Evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
Findings of the 2016 Conference on Machine Translation
Ondrej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno-Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Matt Post, Raphaël Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin M. Verspoor, and Marcos Zampieri. 2016 · 2016
Earlier work this paper cites.
Neural Machine Translation of Rare Words with Subword Units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016 · 2016
Earlier work this paper cites.
The University of Edinburgh’s neural MT systems for WMT17
Rico Sennrich, Alexandra Birch, Anna Currey, Ulrich Germann, Barry Haddow, Kenneth Heafield, Antonio Valerio Miceli Barone, and Philip Williams. 2017 · 2017
Earlier work this paper cites.
Subword Regularization: Improving Neural network Translation Models with Multiple Subword Candidates
Taku Kudo. 2018 · 2018
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing
Taku Kudo and John Richardson. 2018 · 2018
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Earlier work this paper cites.
Investigating the Effectiveness of BPE: The Power of Shorter Sequences
Matthias Gallé. 2019 · 2019
Earlier work this paper cites.
fairseq: A Fast, Extensible Toolkit for Sequence Modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Earlier work this paper cites.
Revisiting low-resource neural machine translation: A case study
Rico Sennrich and Biao Zhang. 2019 · 2019
Earlier work this paper cites.
Byte Pair Encoding is Suboptimal for Language Model Pretraining
Kaj Bostrom and Greg Durrett. 2020 · 2020
Earlier work this paper cites.
BPE-dropout: Simple and effective subword regularization
Ivan Provilkov, Dmitrii Emelianenko, and Elena Voita. 2020 · 2020
Earlier work this paper cites.
COMET: A Neural Framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Earlier work this paper cites.
Superbizarre is not superb: Derivational morphology improves BERT’s interpretation of complex words
Valentin Hofmann, Janet Pierrehumbert, and Hinrich Schütze. 2021 · 2021
Cited alongside, same era.
How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models
Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, and Iryna Gurevych. 2021 · 2021
Cited alongside, same era.
A statistical extension of byte-pair encoding
David Vilar and Marcello Federico. 2021 · 2021
Cited alongside, same era.
The Flores-101 Evaluation Benchmark for Low-Resource and Multilingual Machine Translation
Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan. 2022 · 2022
Cited alongside, same era.
An Embarrassingly Simple Method to Mitigate Undesirable Properties of Pretrained Language Model Tokenizers
Valentin Hofmann, Hinrich Schuetze, and Janet Pierrehumbert. 2022 · 2022
Cited alongside, same era.
An Analysis of BPE Vocabulary Trimming in Neural Machine Translation
Marco Cognetta, Tatsuya Hiraoka, Naoaki Okazaki, Rico Sennrich, and Yuval Pinter. 2024 · 2024
Closest in time.
Coercing LLMs to do and reveal (almost) anything
Jonas Geiping, Alex Stein, Manli Shu, Khalid Saifullah, Yuxin Wen, and Tom Goldstein. 2024 · 2024
Closest in time.
Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance
Omer Goldman, Avi Caciularu, Matan Eyal, Kris Cao, Idan Szpektor, and Reut Tsarfaty. 2024 · 2024
Closest in time.
OLMo: Accelerating the Science of Language Models
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
TextPruner: A Model Pruning Toolkit for Pre-Trained Language Models
Ziqing Yang, Yiming Cui, and Zhigang Chen. 2022 · 2022
Cited alongside, same era.
Tokenizer choice for LLM training: Negligible or Crucial?
Mehdi Ali, Michael Fromm, Klaudia Thellmann, Richard Rutmann, Max Lübbering, Johannes Leveling, Katrin Klug, Jan Ebert, Niclas Doll, Jasper Schulze Buschhoff, et al. 2023 · 2023
Cited alongside, same era.
Analyzing cognitive plausibility of subword tokenization
Lisa Beinborn and Yuval Pinter. 2023 · 2023
Cited alongside, same era.
XLM-V: Overcoming the Vocabulary Bottleneck in Multilingual Masked Language Models
Davis Liang, Hila Gonen, Yuning Mao, Rui Hou, Naman Goyal, Marjan Ghazvininejad, Luke Zettlemoyer, and Madian Khabsa. 2023 · 2023
Cited alongside, same era.
Language Model Tokenizers Introduce Unfairness Between Languages
Aleksandar Petrov, Emanuele La Malfa, Philip Torr, and Adel Bibi. 2023 · 2023
Cited alongside, same era.
Solidgoldmagikarp (plus, prompt generation)
Jessica Rumbelow and Matthew Watkins. 2023 · 2023
Cited alongside, same era.
Impact of tokenization on language models: An analysis for turkish
Cagri Toraman, Eyup Halit Yilmaz, Furkan Şahinuç, and Oguzhan Ozcelik. 2023 · 2023
Cited alongside, same era.
Estonian-Centric Machine Translation: Data, Models, and Challenges
Elizaveta Korotkova and Mark Fishel. 2024 · 2024
Closest in time.
Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models
Sander Land and Max Bartolo. 2024 · 2024
Closest in time.
Glitch tokens in large language models: Categorization taxonomy and effective detection
Yuxi Li, Yi Liu, Gelei Deng, Ying Zhang, Wenjia Song, Ling Shi, Kailong Wang, Yuekang Li, Yang Liu, and Haoyu Wang. 2024 · 2024
Closest in time.
Scaffold-bpe: Enhancing byte pair encoding with simple and effective scaffold token removal
Haoran Lian, Yizhe Xiong, Jianwei Niu, Shasha Mo, Zhenpeng Su, Zijia Lin, Peng Liu, Hui Chen, and Guiguang Ding. 2024 · 2024
Closest in time.
Jiayun Pang and Ivan Vulić. 2024 · 2024
Closest in time.
Toward a Theory of Tokenization in LLMs
Nived Rajaraman, Jiantao Jiao, and Kannan Ramchandran. 2024 · 2024
Closest in time.
Tokenization Is More Than Compression
Craig W Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine, Omri Uzan, Yuval Pinter, and Chris Tanner. 2024 · 2024
Closest in time.
Flexibly Scaling Large Language Models Contexts Through Extensible Tokenization
Ninglu Shao, Shitao Xiao, Zheng Liu, and Peitian Zhang. 2024 · 2024
Closest in time.
Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs
Aaditya K Singh and DJ Strouse. 2024 · 2024
Closest in time.
Tokenization Matters! Degrading Large Language Models through Challenging Their Tokenization
Dixuan Wang, Yanda Li, Junyuan Jiang, Zepeng Ding, Guochao Jiang, Jiaqing Liang, and Deqing Yang. 2024 · 2024
Closest in time.
An Empirical Study on Cross-lingual Vocabulary Adaptation for Efficient Generative LLM Inference
Atsuki Yamaguchi, Aline Villavicencio, and Nikolaos Aletras. 2024 · 2024
Closest in time.
Fast WordPiece Tokenization
Xinying Song, Alex Salcianu, Yang Song, Dave Dopson, and Denny Zhou. 2021 · 2089
Closest in time.