Fetching the paper…
Reading the bibliography…
Tokenization is a foundational step in natural language processing (NLP) tasks, bridging raw text and language models.
Individual comparisons by ranking methods. biom. bull., 1, 80–83
F Wilcoxon. 1945 · 1945
Earlier work this paper cites.
Error bounds for convolutional codes and an asymptotically optimum decoding algorithm
A. Viterbi. 1967 · 1967
Earlier work this paper cites.
Random sampling with a reservoir
Jeffrey S. Vitter. 1985 · 1985
Earlier work this paper cites.
A new algorithm for data compression
Philip Gage. 1994 · 1994
Earlier work this paper cites.
Tokenization , pages 117–133. Springer Netherlands, Dordrecht
Gregory Grefenstette. 1999 · 1999
Earlier work this paper cites.
Normalized (pointwise) mutual information in collocation extraction
Gerlof Bouma. 2009 · 2009
Earlier work this paper cites.
Weighted random sampling over data streams
Pavlos S. Efraimidis. 2010 · 2010
Earlier work this paper cites.
Subword language modeling with neural networks
Tomas Mikolov, Ilya Sutskever, Anoop Deoras, Hai Son Le, Stefan Kombrink, and Jan Honza Černocký. 2011 · 2011
Earlier work this paper cites.
The winograd schema challenge
Hector J. Levesque, Ernest Davis, and Leora Morgenstern. 2012 · 2012
Earlier work this paper cites.
Japanese and korean voice search
Mike Schuster and Kaisuke Nakajima. 2012 · 2012
Earlier work this paper cites.
Qa4mre 2011-2013: Overview of question answering for machine reading evaluation
Anselmo Peñas, Eduard Hovy, Pamela Forner, Álvaro Rodrigo, Richard Sutcliffe, and Roser Morante. 2013 · 2013
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016 · 2016
Earlier work this paper cites.
Race: Large-scale reading comprehension dataset from examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017 · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Johannes Welbl, Nelson F. Liu, and Matt Gardner. 2017 · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018 · 2018
Earlier work this paper cites.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo. 2018 · 2018
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Earlier work this paper cites.
Mathqa: Towards interpretable math word problem solving with operation-based formalisms
Aida Amini, Saadia Gabriel, Peter Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019 · 2019
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
Investigating the effectiveness of BPE: The power of shorter sequences
Matthias Gallé. 2019 · 2019
Cited alongside, same era.
BERT is not an interlingua and the bias of tokenization
Jasdeep Singh, Bryan McCann, Richard Socher, and Caiming Xiong. 2019 · 2019
Cited alongside, same era.
Byte pair encoding is suboptimal for language model pretraining
Kaj Bostrom and Greg Durrett. 2020 · 2020
Cited alongside, same era.
Cassandra L Jacobs and Yuval Pinter. 2022 · 2022
Later among the works it cites.
A framework for few-shot language model evaluation
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2023 · 2023
Later among the works it cites.
The minipile challenge for data-efficient language models
Jean Kaddour. 2023 · 2023
Later among the works it cites.
XLM-V: Overcoming the vocabulary bottleneck in multilingual masked language models
Davis Liang, Hila Gonen, Yuning Mao, Rui Hou, Naman Goyal, Marjan Ghazvininejad, Luke Zettlemoyer, and Madian Khabsa. 2023 · 2023
Later among the works it cites.
Tokenization impacts multilingual language modeling: Assessing vocabulary allocation and overlap across languages
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020 · 2020
Cited alongside, same era.
Dynamic programming encoding for subword segmentation in neural machine translation
Xuanli He, Gholamreza Haffari, and Mohammad Norouzi. 2020 · 2020
Cited alongside, same era.
Getting the ##life out of living: How adequate are word-pieces for modelling complex morphology?
Stav Klein and Reut Tsarfaty. 2020 · 2020
Cited alongside, same era.
BPE-dropout: Simple and effective subword regularization
Ivan Provilkov, Dmitrii Emelianenko, and Elena Voita. 2020 · 2020
Cited alongside, same era.
From characters to words: the turning point of BPE merges
Ximena Gutierrez-Vasques, Christian Bentz, Olga Sozinova, and Tanja Samardzic. 2021 · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Cited alongside, same era.
Superbizarre is not superb: Derivational morphology improves BERT’s interpretation of complex words
Valentin Hofmann, Janet Pierrehumbert, and Hinrich Schütze. 2021 · 2021
Cited alongside, same era.
Tomasz Limisiewicz, Jiří Balhar, and David Mareček. 2023 · 2023
Later among the works it cites.
What changes when you randomly choose BPE merge operations? not much
Jonne Saleva and Constantine Lignos. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 · 2023
Later among the works it cites.
Incorporating context into subword vocabularies
Shaked Yehezkel and Yuval Pinter. 2023 · 2023
Later among the works it cites.
A formal perspective on byte-pair encoding
Vilém Zouhar, Clara Meister, Juan Gastaldi, Li Du, Tim Vieira, Mrinmaya Sachan, and Ryan Cotterell. 2023b · 2023
Later among the works it cites.
Tokenizer choice for llm training: Negligible or crucial?
Mehdi Ali, Michael Fromm, Klaudia Thellmann, Richard Rutmann, Max Lübbering, Johannes Leveling, Katrin Klug, Jan Ebert, Niclas Doll, Jasper Schulze Buschhoff, Charvi Jain, Alexander Arno Weber, Lena Jurkschat, Hammam Abdelwahab, Chelsea John, Pedro Ortiz Suarez, Malte Ostendorff, Samuel Weinbach, Rafet Sifa, Stefan Kesselheim, and Nicolas Flores-Herr. 2024 · 2024
Closest in time.
BPE-knockout: Pruning pre-existing BPE tokenisers with backwards-compatible morphological semi-supervision
Thomas Bauwens and Pieter Delobelle. 2024 · 2024
Closest in time.
Bpe gets picky: Efficient vocabulary refinement during tokenizer training
Pavel Chizhov, Catherine Arnett, Elizaveta Korotkova, and Ivan P. Yamshchikov. 2024 · 2024
Closest in time.
Two counterexamples to tokenization and the noiseless channel
Marco Cognetta, Vilém Zouhar, Sangwhan Moon, and Naoaki Okazaki. 2024 · 2024
Closest in time.
Unpacking tokenization: Evaluating text compression and its correlation with model performance
Omer Goldman, Avi Caciularu, Matan Eyal, Kris Cao, Idan Szpektor, and Reut Tsarfaty. 2024 · 2024
Closest in time.
Word boundary information isn’t useful for encoder language models
Edward Gow-Smith, Dylan Phelps, Harish Tayyar Madabushi, Carolina Scarton, and Aline Villavicencio. 2024 · 2024
Closest in time.
Scaffold-bpe: Enhancing byte pair encoding with simple and effective scaffold token removal
Haoran Lian, Yizhe Xiong, Jianwei Niu, Shasha Mo, Zhenpeng Su, Zijia Lin, Peng Liu, Hui Chen, and Guiguang Ding. 2024 · 2024
Closest in time.
Greed is all you need: An evaluation of tokenizer inference methods
Omri Uzan, Craig W. Schmidt, Chris Tanner, and Yuval Pinter. 2024 · 2024
Closest in time.