Fetching the paper…
Reading the bibliography…
Tokenization imposes a fixed granularity on the input text, freezing how a language model operates on data and how far in the future it predicts.
U-Net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Hierarchical transformers are more efficient language models
Piotr Nawrot, Szymon Tworkowski, Michał Tyrolski, Lukasz Kaiser, Yuhuai Wu, Christian Szegedy, and Henryk Michalewski · 2022
Earlier work this paper cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Oriol Vinyals, Jack W. Rae, and Laurent Sifre · 2022
Earlier work this paper cites.
MEGABYTE: Predicting million-byte sequences with multiscale transformers
Lili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, and Mike Lewis · 2023
Earlier work this paper cites.
Better & faster large language models via multi-token prediction
Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Roziere, David Lopez-Paz, and Gabriel Synnaeve · 2024
Earlier work this paper cites.
Meta Lingua: A minimal PyTorch LLM training library, 2024
Mathurin Videau, Badr Youbi Idrissi, Daniel Haziza, Luca Wehrstedt, Jade Copet, Olivier Teytaud, and David Lopez-Paz · 2024
Earlier work this paper cites.
Byte latent transformer: Patches scale better than tokens
Artidoro Pagnoni, Ram Pasunuru, Pedro Rodriguez, et al · 2024
Earlier work this paper cites.
Getting the most out of your tokenizer for pre-training and domain adaptation
Gautier Dagan, Gabriel Synnaeve, and Baptiste Roziere · 2024
Cited alongside, same era.
SpaceByte: Towards deleting tokenization from large language modeling
Kevin Slagle · 2024
Cited alongside, same era.
Deepseek llm: Scaling open-source language models with longtermism
Xiao Bi, Deli Chen, Guanting Chen, Shanhuang Chen, Damai Dai, Chengqi Deng, Honghui Ding, Kai Dong, Qiushi Du, Zhe Fu, et al · 2024
Cited alongside, same era.
Language models scale reliably with over-training and on downstream tasks
Samir Yitzhak Gadre, Georgios Smyrnis, Vaishaal Shankar, Suchin Gururangan, Mitchell Wortsman, Rulin Shao, Jean Mercat, Alex Fang, Jeffrey Li, Sedrick Keh, et al · 2024
Cited alongside, same era.
DataComp-LM: In search of the next generation of training sets for language models
Jeffrey Li, Alex Fang, Georgios Smyrnis, et al · 2024
Tokenizer choice for LLM training: Negligible or crucial?
Mehdi Ali, Michael Fromm, Klaudia Thellmann, et al · 2024
Later among the works it cites.
An analysis of tokenization: Transformers under markov data
Nived Rajaraman, Jiantao Jiao, and Kannan Ramchandran · 2024
Later among the works it cites.
Retok: Replacing tokenizer to enhance representation efficiency in large language model
Shuhao Gu, Mengdi Zhao, Bowen Zhang, Liangdong Wang, Jijie Li, and Guang Liu · 2024
Later among the works it cites.
Training LLMs over neurally compressed text
Brian Lester, Jaehoon Lee, Alexander A Alemi, Jeffrey Pennington, Adam Roberts, Jascha Sohl-Dickstein, and Noah Constant · 2024
Later among the works it cites.
Enhancing large language models through adaptive tokenizers
Mengyu Zheng, Hanting Chen, Tianyu Guo, Chong Zhu, Binfan Zheng, Chang Xu, and Yunhe Wang · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2024
Cited alongside, same era.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, et al · 2024
Cited alongside, same era.
CUTE: Measuring LLMs’ understanding of their tokens
Lukas Edman, Helmut Schmid, and Alexander Fraser · 2024
Cited alongside, same era.
Scaling neural machine translation to 200 languages
Marta Costa-jussa, James Cross, Onur Çelebi, Maha Elbayad, et al · 2024
Cited alongside, same era.
T-FREE: Subword tokenizer-free generative LLMs via sparse representations for memory-efficient embeddings
Björn Deiseroth, Manuel Brack, Patrick Schramowski, Kristian Kersting, and Samuel Weinbach · 2024
Later among the works it cites.
Vincent-Pierre Berges, Barlas Oğuz, Daniel Haziza, Wen-tau Yih, Luke Zettlemoyer, and Gargi Ghosh · 2024
Later among the works it cites.
Hierarchical autoregressive transformers: Combining byte- and word-level processing for robust, adaptable language models
Pit Neitemeier, Björn Deiseroth, Constantin Eichenberg, and Lukas Balles · 2025
Closest in time.