Fetching the paper…
Reading the bibliography…
We demonstrate the existence of a complexity phase transition in neural networks by studying the grokking phenomenon, where networks suddenly transition from memorization to generalization long after overfitting their training data.
A mathematical theory of communication,
C. E. Shannon, · 1948
Earlier work this paper cites.
On the uniform convergence of relative frequencies of events to their probabilities,
V. N. Vapnik, A. Chervonenkis, · 1971
Earlier work this paper cites.
Modeling by shortest data description*,
J. Rissanen, · 1978
Earlier work this paper cites.
Principles of risk minimization for learning theory,
V. N. Vapnik, · 1991
Earlier work this paper cites.
An introduction to kolmogorov complexity and its applications,
M. Li, P. M. B. Vitányi, · 1993
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights,
G. E. Hinton, D. van Camp, · 1993
Earlier work this paper cites.
2001
Earlier work this paper cites.
Bounds for averaging classifiers.,
J. Langford, M. Seeger, · 2001
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results,
P. L. Bartlett, S. Mendelson, · 2003
Earlier work this paper cites.
Rate distortion and denoising of individual data using kolmogorov complexity,
N. K. Vereshchagin, P. M. B. Vitányi, · 2004
Earlier work this paper cites.
The effective rank: A measure of effective dimensionality,
O. Roy, M. Vetterli, · 2007
Cited alongside, same era.
Algorithmic thermodynamics,
J. Baez, M. Stay, · 2012
Cited alongside, same era.
2014
Cited alongside, same era.
P. Micikevicius, S. Narang, J. Alben, G. F. Diamos, E. Elsen, D. García, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, H. Wu, · 2017
Cited alongside, same era.
Measuring the intrinsic dimension of objective landscapes,
C. Li, H. Farkhoor, R. Liu, J. Yosinski, · 2018
Cited alongside, same era.
Lora: Low-rank adaptation of large language models,
J. E. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, W. Chen, · 2021
Later among the works it cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets,
A. Power, Y. Burda, H. Edwards, I. Babuschkin, V. Misra, · 2022
Later among the works it cites.
Explaining grokking through circuit efficiency,
V. Varma, R. Shah, Z. Kenton, J. Kram’ar, R. Kumar, · 2023
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability,
N. Nanda, L. Chan, T. Lieberum, J. Smith, J. Steinhardt, · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep double descent: where bigger models and more data hurt,
P. Nakkiran, G. Kaplun, Y. Bansal, T. Yang, B. Barak, I. Sutskever, · 2019
Cited alongside, same era.
Rethinking lossy compression: The rate-distortion-perception tradeoff,
Y. Blau, T. Michaeli, · 2019
Cited alongside, same era.
Scaling laws for autoregressive generative modeling,
T. Henighan, J. Kaplan, M. Katz, M. Chen, C. Hesse, J. Jackson, H. Jun, T. B. Brown, P. Dhariwal, S. Gray, C. Hallacy, B. Mann, A. Radford, A. Ramesh, N. Ryder, D. M. Ziegler, J. Schulman, D. Amodei, S. McCandlish, · 2020
Cited alongside, same era.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning,
A. Aghajanyan, L. Zettlemoyer, S. Gupta, · 2020
Cited alongside, same era.
Thermodynamic costs of turing machines,
A. Kolchinsky, D. H. Wolpert, · 2020
Cited alongside, same era.
Towards understanding grokking: An effective theory of representation learning,
Z. Liu, O. Kitouni, N. Nolte, E. J. Michaud, M. Tegmark, M. Williams,
Cited in the paper.
Omnigrok: Grokking beyond algorithmic data,
Z. Liu, E. J. Michaud, M. Tegmark,
Cited in the paper.
M. Goldblum, M. Finzi, K. Rowan, A. G. Wilson, · 2023
Later among the works it cites.
Qlora: Efficient finetuning of quantized llms,
T. Dettmers, A. Pagnoni, A. Holtzman, L. Zettlemoyer, · 2023
Later among the works it cites.
Non-vacuous generalization bounds for large language models,
S. Lotfi, M. Finzi, Y. Kuang, T. G. J. Rudner, M. Goldblum, A. G. Wilson, · 2024
Closest in time.
The era of 1-bit llms: All large language models are in 1.58 bits,
S. Ma, H. Wang, L. Ma, L. Wang, W. Wang, S. Huang, L. Dong, R. Wang, J. Xue, F. Wei, · 2024
Closest in time.
Foundations of algorithmic thermodynamics,
A. Ebtekar, M. Hutter, · 2025
Closest in time.