Fetching the paper…
Reading the bibliography…
In 2020, OpenAI proposed the first type of Scaling Laws, describing the relationships between model loss and the scale of parameters, data, and training computation.
On computable numbers, with an application to the entscheidungs problem
A Turing. 1936 · 1936
Earlier work this paper cites.
The complexity of finite objects and the development of the concepts of information and randomness by means of the theory of algorithms
Alexander K Zvonkin and Leonid A Levin. 1970 · 1970
Earlier work this paper cites.
Source coding algorithms for fast data compression
Richard Clark Pasco. 1976 · 1976
Earlier work this paper cites.
Generalized kraft inequality and arithmetic coding
Jorma J Rissanen. 1976 · 1976
Earlier work this paper cites.
Modeling by shortest data description
Jorma Rissanen. 1978 · 1978
Earlier work this paper cites.
A universal prior for integers and estimation by minimum description length
Jorma Rissanen. 1983 · 1983
Earlier work this paper cites.
Kolmogorov Complexity, Data Compression, and Inference , pages 23–33
Thomas M. Cover. 1985 · 1985
Earlier work this paper cites.
Stochastic complexity and modeling
Jorma Rissanen. 1986 · 1986
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, and Pascal Vincent. 2000 · 2000
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
50,000 euro prize for compressing human knowledge
Marcus Hutter. 2006 · 2006
Cited alongside, same era.
An introduction to Kolmogorov complexity and its applications , volume 3
Ming Li, Paul Vitányi, et al. 2008 · 2008
Cited alongside, same era.
A machine learning perspective on predictive coding with paq8
Byron Knoll and Nando de Freitas. 2012 · 2012
Cited alongside, same era.
Nncp: Lossless data compression with neural networks
F Bellard. 2019 · 2019
Cited alongside, same era.
Nncp v2: Lossless data compression with transformer
Fabrice Bellard. 2021 · 2021
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. 2022 · 2022
Cited alongside, same era.
The complexity dynamics of grokking
Branton DeMoss, Silvia Sapora, Jakob Foerster, Nick Hawes, and Ingmar Posner. 2024 · 2024
Later among the works it cites.
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. 2024 · 2024
Later among the works it cites.
How powerful are decoder-only transformer neural models?
Jesse Roberts. 2024 · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. 2024 · 2024
Later among the works it cites.
Building effective agents
Anthropic. 2024 · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Language modeling is compression
Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, et al. 2023 · 2023
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao. 2023 · 2023
Cited alongside, same era.
Rwkv: Reinventing rnns for the transformer era
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Stella Biderman, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, et al. 2023 · 2023
Cited alongside, same era.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. 2025 · 2025
Closest in time.
Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori Hashimoto. 2025 · 2025
Closest in time.
Learning to reason with llms
OpenAI. 2024 · 2025
Closest in time.
entropix
xjdr-alt. 2024 · 2025
Closest in time.