Fetching the paper…
Reading the bibliography…
In this paper, we explore the idea of training large language models (LLMs) over highly compressed text.
Generating Long Sequences with Sparse Transformers, April 2019
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 1904
Earlier work this paper cites.
Bridging the Gap for Tokenizer-Free Language Models
Dokook Choe, Rami Al-Rfou, Mandy Guo, Heeyoung Lee, and Noah Constant · 1908
Earlier work this paper cites.
Does Knowledge Distillation Really Work?
Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alex Alemi, and Andrew Gordon Wilson · 1911
Earlier work this paper cites.
The Psycho-Biology of Language
G. K. Zipf · 1935
Earlier work this paper cites.
The Generalization of “Student’s” Problem when Several Different Population Variances are Involved
Bernard Lewis Welch · 1947
Earlier work this paper cites.
A Mathematical Theory of Communication
Claude Elwood Shannon · 1948
Earlier work this paper cites.
On Information and Sufficiency
S. Kullback and R. A. Leibler · 1951
Earlier work this paper cites.
A Method for the Construction of Minimum-Redundancy Codes
David A. Huffman · 1952
Earlier work this paper cites.
Note on the Bias of Information Estimates
George Miller · 1955
Earlier work this paper cites.
Binary codes capable of correcting deletions, insertions, and reversals
Vladimir I. Levenshtein · 1965
Earlier work this paper cites.
Cognitron: A Self-Organizing Multilayered Neural Network
Kunihiko Fukushima · 1975
Earlier work this paper cites.
Generalized Kraft Inequality and Arithmetic Coding
J. J. Rissanen · 1976
Earlier work this paper cites.
Source Coding Algorithms for Fast Data Compression (Ph.D. Thesis Abstr.)
R. Pasco · 1977
Earlier work this paper cites.
A Universal Algorithm for Sequential Data Compression
J. Ziv and A. Lempel · 1977
Earlier work this paper cites.
Bootstrap Methods: Another Look at the Jackknife
B. Efron · 1979
Earlier work this paper cites.
Arithmetic Coding for Data Compression
Ian H. Witten, Radford M. Neal, and John G. Cleary · 1987
Earlier work this paper cites.
Analysis of Arithmetic Coding for Data Compression
Paul G. Howard and Jeffrey Scott Vitter · 1992
Earlier work this paper cites.
A New Algorithm for Data Compression
Philip Gage · 1994
Earlier work this paper cites.
GZIP file format specification
P Deutsch · 1996
Earlier work this paper cites.
ZLIB Compressed Data Format Specification
P Deutsch and J-L Gailly · 1996
Earlier work this paper cites.
Greedy Function Approximation: A Gradient Boosting Machine
Jerome H. Friedman · 2001
Earlier work this paper cites.
Scaling Laws for Neural Language Models, January 2020
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2001
Earlier work this paper cites.
The Mathematics of Statistical Machine Translation: Parameter Estimation
Peter F. Brown, Vincent J. Della Pietra, Stephen A. Della Pietra, and Robert L. Mercer · 2003
Earlier work this paper cites.
Estimation of Entropy and Mutual Information
Liam Paninski · 2003
Earlier work this paper cites.
Longformer: The Long-Document Transformer, December 2020
Iz Beltagy, Matthew E. Peters, and Arman Cohan · 2004
Earlier work this paper cites.
Matplotlib: A 2D Graphics Environment
J. D. Hunter · 2007
Earlier work this paper cites.
The Complexity of Finite Objects and the Development of the Concepts of Information and Randomness by Means of the Theory of Algorithms
A Zvonkin and L Levin · 2007
Earlier work this paper cites.
Python 3 Reference Manual
Guido Van Rossum and Fred L. Drake · 2009
Earlier work this paper cites.
Rectified Linear Units Improve Restricted Boltzmann Machines
Vinod Nair and Geoffrey E. Hinton · 2010
Earlier work this paper cites.
Compressed Context Modeling for Text Compression
M Oğuzhan Külekci · 2011
Cited alongside, same era.
SentencePiece: A Simple and Language Independent Subword Tokenizer and Detokenizer for Neural Text Processing
Taku Kudo and John Richardson · 2012
Cited alongside, same era.
Japanese and Korean Voice Search
Mike Schuster and Kaisuke Nakajima · 2012
Cited alongside, same era.
Bootstrap Methods for the Empirical Study of Decision-Making and Information Flows in Social Systems
Simon DeDeo, Robert X. D. Hawkins, Sara Klingenstein, and Tim Hitchcock · 2013
Cited alongside, same era.
Data Compression Explained
Matt Mahoney · 2013
Cited alongside, same era.
Are Transformers Universal Approximators of Sequence-to-Sequence Functions?
Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank Reddi, and Sanjiv Kumar · 2020
Later among the works it cites.
Big Bird: Transformers for Longer Sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al · 2020
Later among the works it cites.
Yu-An Chung, Yu Zhang, Wei Han, Chung-Cheng Chiu, James Qin, Ruoming Pang, and Yonghui Wu · 2021
Later among the works it cites.
Sequence Length is a Domain: Length-based Overfitting in Transformer Models
Dusan Varis and Ondřej Bojar · 2021
Later among the works it cites.
Seaborn: Statistical Data Visualization
Michael L. Waskom · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jarek Duda · 2014
Cited alongside, same era.
Distilling the Knowledge in a Neural Network, March 2015
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Cited alongside, same era.
Gorilla: a Fast, Scalable, In-Memory Time Series Database
Tuomas Pelkonen, Scott Franklin, Justin Teller, Paul Cavallaro, Qi Huang, Justin Meza, and Kaushik Veeraraghavan · 2015
Cited alongside, same era.
Neural Machine Translation of Rare Words with Subword Units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Cited alongside, same era.
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean · 2016
Cited alongside, same era.
Hierarchical Multiscale Recurrent Neural Networks
Junyoung Chung, Sungjin Ahn, and Yoshua Bengio · 2017
Cited alongside, same era.
Adaptive Computation Time for Recurrent Neural Networks, February 2017
Alex Graves · 2017
Cited alongside, same era.
Canine: Pre-training an Efficient Tokenization-Free Encoder for Language Representation
Jonathan H. Clark, Dan Garrette, Iulia Turc, and John Wieting · 2022
Later among the works it cites.
MANTa: Efficient Gradient-Based Tokenization for End-to-End Robust Language Modeling
Nathan Godey, Roman Castagné, Éric de la Clergerie, and Benoît Sagot · 2022
Later among the works it cites.
Training Compute-Optimal Large Language Models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katherine Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Oriol Vinyals, Jack William Rae, and Laurent Sifre · 2022
Later among the works it cites.
Less is More: Parameter-Free Text Classification with Gzip, December 2022
Zhiying Jiang, Matthew Y. R. Yang, Mikhail Tsirlin, Raphael Tang, and Jimmy Lin · 2022
Later among the works it cites.
Reducing Retraining by Recycling Parameter-Efficient Prompts
Brian Lester, Joshua Yurtsever, Siamak Shakeri, and Noah Constant · 2022
Later among the works it cites.
Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
Ofir Press, Noah Smith, and Mike Lewis · 2022
Later among the works it cites.
Charformer: Fast Character Transformers via Gradient-based Subword Tokenization
Yi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu, and Donald Metzler · 2022
Later among the works it cites.
ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models
Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel · 2022
Later among the works it cites.
AudioLM: A Language Modeling Approach to Audio Generation
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, and Neil Zeghidour · 2023
Later among the works it cites.
Fast Inference from Transformers via Speculative Decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias · 2023
Later among the works it cites.
Character-Aware Models Improve Visual Text Rendering
Rosanne Liu, Dan Garrette, Chitwan Saharia, William Chan, Adam Roberts, Sharan Narang, Irina Blok, Rj Mical, Mohammad Norouzi, and Noah Constant · 2023
Later among the works it cites.
Scaling Up Models and Data with t5x
Adam Roberts, Hyung Won Chung, Gaurav Mishra, Anselm Levskaya, James Bradbury, Daniel Andor, Sharan Narang, Brian Lester, Colin Gaffney, Afroz Mohiuddin, Curtis Hawthorne, Aitor Lewkowycz, Alex Salcianu, Marc van Zee, Jacob Austin, Sebastian Goodman, Livio Baldini Soares, Haitang Hu, Sasha Tsvyashchenko, Aakanksha Chowdhery, Jasmijn Bastings, Jannis Bulian, Xavier Garcia, Jianmo Ni, Andrew Chen, Kathleen Kenealy, Kehang Han, Michelle Casbon, Jonathan H. Clark, Stephan Lee, Dan Garrette, James Lee-Thorp, Colin Raffel, Noam Shazeer, Marvin Ritter, Maarten Bosma, Alexandre Passos, Jeremy Maitin-Shepard, Noah Fiedel, Mark Omernick, Brennan Saeta, Ryan Sepassi, Alexander Spiridonov, Joshua Newlan, and Andrea Gesmundo · 2023
Later among the works it cites.
Are Emergent Abilities of Large Language Models a Mirage?
Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo · 2023
Later among the works it cites.
LLMZip: Lossless Text Compression using Large Language Models, June 2023
Chandra Shekhara Kaushik Valmeekam, Krishna Narayanan, Dileep Kalathil, Jean-Francois Chamberland, and Srinivas Shakkottai · 2023
Later among the works it cites.
Arithmetic Sampling: Parallel Diverse Decoding for Large Language Models
Luke Vilnis, Yury Zemlyanskiy, Patrick Murray, Alexandre Tachard Passos, and Sumit Sanghai · 2023
Later among the works it cites.
Tokenizer choice for LLM training: Negligible or crucial?
Mehdi Ali, Michael Fromm, Klaudia Thellmann, Richard Rutmann, Max Lübbering, Johannes Leveling, Katrin Klug, Jan Ebert, Niclas Doll, Jasper Buschhoff, Charvi Jain, Alexander Weber, Lena Jurkschat, Hammam Abdelwahab, Chelsea John, Pedro Ortiz Suarez, Malte Ostendorff, Samuel Weinbach, Rafet Sifa, Stefan Kesselheim, and Nicolas Flores-Herr · 2024
Closest in time.
TensorFlow Compression: Learned Data Compression, 2024
Johannes Ballé, Sung Jin Hwang, and Eirikur Agustsson · 2024
Closest in time.
Getting the Most Out of Your Tokenizer for Pre-training and Domain Adaptation
Gautier Dagan, Gabriel Synnaeve, and Baptiste Rozière · 2024
Closest in time.
Language Modeling Is Compression
Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, Marcus Hutter, and Joel Veness · 2024
Closest in time.
Omer Goldman, Avi Caciularu, Matan Eyal, Kris Cao, Idan Szpektor, and Reut Tsarfaty · 2024
Closest in time.
Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains
Ashok Vardhan Makkuva, Marco Bondaschi, Adway Girish, Alliot Nagle, Martin Jaggi, Hyeji Kim, and Michael Gastpar · 2024
Closest in time.
Toward a Theory of Tokenization in LLMs
Nived Rajaraman, Jiantao Jiao, and Kannan Ramchandran · 2024
Closest in time.
Tokenization is more than compression, 2024
Craig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine, Omri Uzan, Yuval Pinter, and Chris Tanner · 2024
Closest in time.
Self-Attention with Relative Position Representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2074
Closest in time.