Fetching the paper…
Reading the bibliography…
It has long been established that predictive models can be transformed into lossless compressors and vice versa.
A mathematical theory of communication
Claude E. Shannon · 1948
Earlier work this paper cites.
A method for the construction of minimum-redundancy codes
David A. Huffman · 1952
Earlier work this paper cites.
Generalized kraft inequality and arithmetic coding
Jorma Rissanen · 1976
Earlier work this paper cites.
Source coding algorithms for fast data compression (ph.d. thesis abstr.)
Richard C. Pasco · 1977
Earlier work this paper cites.
Data compression using adaptive coding and partial string matching
John G. Cleary and Ian H. Witten · 1984
Earlier work this paper cites.
A technique for high-performance data compression
Terry A. Welch · 1984
Earlier work this paper cites.
Arithmetic coding for data compression
Ian H. Witten, Radford M. Neal, and John G. Cleary · 1987
Earlier work this paper cites.
Analysis of arithmetic coding for data compression
Paul G. Howard and Jeffrey Scott Vitter · 1991
Earlier work this paper cites.
Predictive coding with neural nets: Application to text compression
Jürgen Schmidhuber and Stefan Heil · 1994
Earlier work this paper cites.
The context-tree weighting method: basic properties
Frans M. J. Willems, Yuri M. Shtarkov, and Tjalling J. Tjalkens · 1995
Earlier work this paper cites.
GZIP file format specification version 4.3
Peter Deutsch · 1996
Earlier work this paper cites.
Sequential neural text compression
Jürgen Schmidhuber and Stefan Heil · 1996
Earlier work this paper cites.
PNG (portable network graphics) specification version 1.0
Thomas Boutell · 1997
Earlier work this paper cites.
On tables of random numbers
Andrei N. Kolmogorov · 1998
Earlier work this paper cites.
Text categorization using compression models
Eibe Frank, Chang Chui, and Ian H. Witten · 2000
Earlier work this paper cites.
Fast text compression with neural networks
Matthew V. Mahoney · 2000
Earlier work this paper cites.
Information theory, inference, and learning algorithms
David J. C. MacKay · 2003
Earlier work this paper cites.
Using Compression-Based Language Models for Text Categorization , pp. 141–165
William J. Teahan and David J. Harper · 2003
Earlier work this paper cites.
Universal Artificial Intellegence - Sequential Decisions Based on Algorithmic Probability
Marcus Hutter · 2005
Earlier work this paper cites.
500’000€ prize for compressing human knowledge, 2006
Marcus Hutter · 2006
Earlier work this paper cites.
Free lossless audio codec, 2008
Josh Coalson · 2008
Earlier work this paper cites.
Jarek Duda · 2009
Earlier work this paper cites.
A philosophical treatise of universal induction
Samuel Rathmanner and Marcus Hutter · 2011
Earlier work this paper cites.
A machine learning perspective on predictive coding with PAQ8
Byron Knoll and Nando de Freitas · 2012
Earlier work this paper cites.
Statistical Language Models Based on Neural Networks
Tomas Mikolov · 2012
Earlier work this paper cites.
CMIX, 2014
Byron Knoll · 2014
Earlier work this paper cites.
The student-t mixture as a natural image patch prior with application to image compression
Aäron van den Oord and Benjamin Schrauwen · 2014
Cited alongside, same era.
Librispeech: An ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael S. Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Cited alongside, same era.
Compress and control
Joel Veness, Marc G. Bellemare, Marcus Hutter, Alvin Chua, and Guillaume Desjardins · 2015
Cited alongside, same era.
Syntactically informed text compression with recurrent neural networks
David Cox · 2016
Cited alongside, same era.
Learning better lossless compression using lossy compression
Fabian Mentzer, Luc Van Gool, and Michael Tschannen · 2020
Later among the works it cites.
Bpe-dropout: Simple and effective subword regularization
Ivan Provilkov, Dmitrii Emelianenko, and Elena Voita · 2020
Later among the works it cites.
Deep-learning-based lossless image coding
Ionut Schiopu and Adrian Munteanu · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontañón, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed · 2020
Later among the works it cites.
NNCP v2: Lossless data compression with transformer
Fabrice Bellard · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Cited alongside, same era.
A survey of model compression and acceleration for deep neural networks
Yu Cheng, Duo Wang, Pan Zhou, and Tao Zhang · 2017
Cited alongside, same era.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
The description length of deep learning models
Léonard Blier and Yann Ollivier · 2018
Cited alongside, same era.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo · 2018
Cited alongside, same era.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson · 2018
Cited alongside, same era.
Rishi Bommasani et al · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
Jack W. Rae et al · 2021
Later among the works it cites.
Accelerated deep lossless image coding with unified paralleleized GPU coding architecture
Benjamin Lukas Cajus Barzen, Fedor Glazov, Jonas Geistert, and Thomas Sikora · 2022
Later among the works it cites.
Longt5: Efficient text-to-text transformer for long sequences
Mandy Guo, Joshua Ainslie, David C. Uthus, Santiago Ontañón, Jianmo Ni, Yun-Hsuan Sung, and Yinfei Yang · 2022
Later among the works it cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, et al · 2022
Later among the works it cites.
Few-shot non-parametric learning with deep latent variable model
Zhiying Jiang, Yiqin Dai, Ji Xin, Ming Li, and Jimmy Lin · 2022
Later among the works it cites.
TRACE: A fast transformer-based general-purpose lossless compressor
Yu Mao, Yufei Cui, Tei-Wei Kuo, and Chun Jason Xue · 2022
Later among the works it cites.
LC-FDNet: Learned lossless image compression with frequency decomposition network
Hochang Rhee, Yeong Il Jang, Seyun Kim, and Nam Ik Cho · 2022
Later among the works it cites.
Compression of generative pre-trained language models via quantization
Chaofan Tao, Lu Hou, Wei Zhang, Lifeng Shang, Xin Jiang, Qun Liu, Ping Luo, and Ngai Wong · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou · 2022
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with GPT-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott M. Lundberg, Harsha Nori, Hamid Palangi, Marco Túlio Ribeiro, and Yi Zhang · 2023
Closest in time.
Scaling transformer to 1m tokens and beyond with RMT
Aydar Bulatov, Yuri Kuratov, and Mikhail S. Burtsev · 2023
Closest in time.
Neural networks and the chomsky hierarchy
Grégoire Delétang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein, Li Kevin Wenliang, Elliot Catt, Chris Cundy, Marcus Hutter, Shane Legg, Joel Veness, and Pedro A. Ortega · 2023
Closest in time.
In-context autoencoder for context compression in a large language model
Tao Ge, Jing Hu, Xun Wang, Si-Qing Chen, and Furu Wei · 2023
Closest in time.
Memory-based meta-learning on non-stationary distributions
Tim Genewein, Grégoire Delétang, Anian Ruoss, Li Kevin Wenliang, Elliot Catt, Vincent Dutordoir, Jordi Grau-Moya, Laurent Orseau, Marcus Hutter, and Joel Veness · 2023
Closest in time.
"low-resource" text classification: A parameter-free classification method with compressors
Zhiying Jiang, Matthew Y. R. Yang, Mikhail Tsirlin, Raphael Tang, Yiqin Dai, and Jimmy Lin · 2023
Closest in time.
In-context reinforcement learning with algorithm distillation
Michael Laskin, Luyu Wang, et al · 2023
Closest in time.
Large language models as general pattern machines
Suvir Mirchandani, Fei Xia, Pete Florence, Brian Ichter, Danny Driess, Montserrat Gonzalez Arenas, Kanishka Rao, Dorsa Sadigh, and Andy Zeng · 2023
Closest in time.
Gzip versus bag-of-words for text classification, 2023
Juri Opitz · 2023
Closest in time.
Randomized positional encodings boost length generalization of transformers
Anian Ruoss, Grégoire Delétang, Tim Genewein, Jordi Grau-Moya, Róbert Csordás, Mehdi Bennani, Shane Legg, and Joel Veness · 2023
Closest in time.
Llmzip: Lossless text compression using large language models
Chandra Shekhara Kaushik Valmeekam, Krishna Narayanan, Dileep Kalathil, Jean-François Chamberland, and Srinivas Shakkottai · 2023
Closest in time.