Fetching the paper…
Reading the bibliography…
While there has been a large body of research attempting to circumvent tokenization for language modeling (Clark et al., 2022; Xue et al., 2022), the current consensus is that it is a necessary initial step for designing state-of-the-art performant language models.
An algorithm for the segmentation of an artificial language analogue
J Gerard Wolff · 1975
Earlier work this paper cites.
Compression of individual sequences via variable-rate coding
Jacob Ziv and Abraham Lempel · 1978
Earlier work this paper cites.
A technique for high-performance data compression
Terry A. Welch · 1984
Earlier work this paper cites.
A new algorithm for data compression
Philip Gage · 1994
Earlier work this paper cites.
What is a word, what is a sentence?: problems of tokenisation
Gregory Grefenstette and Pasi Tapanainen · 1994
Earlier work this paper cites.
Off-line dictionary-based compression
N Jesper Larsson and Alistair Moffat · 2000
Earlier work this paper cites.
Tokenisation and sentence segmentation
David D Palmer · 2000
Earlier work this paper cites.
Average profile of the lempel-ziv parsing scheme for a markovian source
Philippe Jacquet, Wojciech Szpankowski, and Jing Tang · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Bernstein polynomials and learning theory
Dietrich Braess and Thomas Sauer · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
The development of an electronic dictionary for morphological analysis and its application to japanese corpus linguistics, Oct 2007
Yasuharu Den, Toshinobu Ogiso, Hideki Ogura, Atsushi Yamada, Nobuaki Minematsu, Kiyotaka Uchimoto, and Hanae Koiso · 2007
Earlier work this paper cites.
Re-pair achieves high-order entropy
Gonzalo Navarro and Luís MS Russo · 2008
Earlier work this paper cites.
Probability, random processes, and ergodic properties , volume 1
Robert M Gray and RM Gray · 2009
Earlier work this paper cites.
Juman++: A morphological analysis toolkit for scriptio continua
Arseny Tolmachev, Daisuke Kawahara, and Sadao Kurohashi · 2010
Earlier work this paper cites.
The expected profile of digital search trees
Michael Drmota and Wojciech Szpankowski · 2011
Earlier work this paper cites.
Japanese and korean voice search
Mike Schuster and Kaisuke Nakajima · 2012
Earlier work this paper cites.
The Definitive ANTLR 4 Reference
Terence Parr · 2013
Earlier work this paper cites.
Typical depth of a digital search tree built on a general source
Kanal Hun and Brigitte Vallée · 2014
Earlier work this paper cites.
Operator theoretic aspects of ergodic theory , volume 272
Tanja Eisner, Bálint Farkas, Markus Haase, and Rainer Nagel · 2015
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Cited alongside, same era.
Stochastic modeling of scientific data, Autumn 2018
Yen-Chi Chen · 2018
Cited alongside, same era.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo · 2018
Cited alongside, same era.
An improper estimator with optimal excess risk in misspecified density estimation and logistic regression
Jaouad Mourtada and Stéphane Gaïffas · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, et al · 2022
Later among the works it cites.
Byt5: Towards a token-free future with pre-trained byte-to-byte models
Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel · 2022
Later among the works it cites.
Tokenizer choice for llm training: Negligible or crucial?
Mehdi Ali, Michael Fromm, Klaudia Thellmann, Richard Rutmann, Max Lübbering, Johannes Leveling, Katrin Klug, Jan Ebert, Niclas Doll, Jasper Schulze Buschhoff, et al · 2023
Later among the works it cites.
Evaluating various tokenizers for arabic text classification
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Matthias Gallé · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Cited alongside, same era.
Language models are few-shot learners
Ben Mann, N Ryder, M Subbiah, J Kaplan, P Dhariwal, A Neelakantan, P Shyam, G Sastry, A Askell, S Agarwal, et al · 2020
Cited alongside, same era.
Concentration of markov chains with bounded moments
Assaf Naor, Shravas Rao, and Oded Regev · 2020
Cited alongside, same era.
Node profiles of symmetric digital search trees: Concentration properties
Michael Drmota, Michael Fuchs, Hsien-Kuei Hwang, and Ralph Neininger · 2021
Cited alongside, same era.
Optimal prediction of markov chains with and without spectral gap
Yanjun Han, Soham Jana, and Yihong Wu · 2021
Cited alongside, same era.
Zaid Alyafeai, Maged S Al-shaibani, Mustafa Ghaleb, and Irfan Ahmad · 2023
Later among the works it cites.
xval: A continuous number encoding for large language models
Siavash Golkar, Mariel Pettee, Michael Eickenberg, Alberto Bietti, Miles Cranmer, Geraud Krawezik, Francois Lanusse, Michael McCabe, Ruben Ohana, Liam Parker, et al · 2023
Later among the works it cites.
Language model tokenizers introduce unfairness between languages
Aleksandar Petrov, Emanuele La Malfa, Philip HS Torr, and Adel Bibi · 2023
Later among the works it cites.
Solidgoldmagikarp
Jessica Rumbelow and Matthew Watkins · 2023
Later among the works it cites.
Impact of tokenization on language models: An analysis for turkish
Cagri Toraman, Eyup Halit Yilmaz, Furkan Şahinuç, and Oguzhan Ozcelik · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Bloomberggpt: A large language model for finance
Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann · 2023
Later among the works it cites.
Outline, then details: Syntactically guided coarse-to-fine code generation
Wenqing Zheng, SP Sharan, Ajay Kumar Jaiswal, Kevin Wang, Yihan Xi, Dejia Xu, and Zhangyang Wang · 2023
Later among the works it cites.
Tokenization and the noiseless channel
Vilém Zouhar, Clara Meister, Juan Gastaldi, Li Du, Mrinmaya Sachan, and Ryan Cotterell · 2023
Later among the works it cites.
The evolution of statistical induction heads: In-context learning markov chains
Benjamin L Edelman, Ezra Edelman, Surbhi Goel, Eran Malach, and Nikolaos Tsilivis · 2024
Closest in time.
Attention with markov: A framework for principled analysis of transformers via markov chains
Ashok Vardhan Makkuva, Marco Bondaschi, Adway Girish, Alliot Nagle, Martin Jaggi, Hyeji Kim, and Michael Gastpar · 2024
Closest in time.
How transformers learn causal structure with gradient descent
Eshaan Nichani, Alex Damian, and Jason D Lee · 2024
Closest in time.
An analysis of tokenization: Transformers under markov data
Nived Rajaraman, Jiantao Jiao, and Kannan Ramchandran · 2024
Closest in time.
Tokenization counts: the impact of tokenization on arithmetic in frontier llms
Aaditya K Singh and DJ Strouse · 2024
Closest in time.