Fetching the paper…
Reading the bibliography…
In some settings neural networks exhibit a phenomenon known as \textit{grokking}, where they achieve perfect or near-perfect accuracy on the validation set long after the same performance has been achieved on the training set.
A mathematical theory of communication
Claude Elwood Shannon · 1948
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
Geoffrey E. Hinton and Drew van Camp · 1993
Earlier work this paper cites.
Minimum description length induction, Bayesianism, and Kolmogorov complexity
Paul MB Vitányi and Ming Li · 2000
Earlier work this paper cites.
Information Theory, Inference, and Learning Algorithms
David J. C. MacKay · 2003
Earlier work this paper cites.
Model Selection and Multimodel Inference: A Practical Information-Theoretic Approach
Kenneth P. Burnham and David R. Anderson · 2004
Earlier work this paper cites.
Elements of Information Theory
Thomas M. Cover and Joy A. Thomas · 2006
Earlier work this paper cites.
Gaussian processes for machine learning
Carl Edward Rasmussen and Christopher K. I. Williams · 2006
Earlier work this paper cites.
Scalable variational Gaussian process classification
James Hensman, Alexander G. de G. Matthews, and Zoubin Ghahramani · 2015
Earlier work this paper cites.
Understanding probabilistic sparse Gaussian process approximations
Matthias Bauer, Mark van der Wilk, and Carl Edward Rasmussen · 2016
Earlier work this paper cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
Gpytorch: Blackbox matrix-matrix Gaussian process inference with GPU acceleration
Jacob Gardner, Geoff Pleiss, Kilian Q Weinberger, David Bindel, and Andrew G Wilson · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Cited alongside, same era.
A Kolmogorov complexity approach to generalization in deep learning
Hazar Yueksel, Kush R Varshney, and Brian Kingsbury · 2019
Cited alongside, same era.
How incomputable is Kolmogorov complexity?
Paul MB Vitányi · 2020
Cited alongside, same era.
Model complexity of deep learning: A survey
Xia Hu, Lingyang Chu, Jian Pei, Weiqing Liu, and Jiang Bian · 2021
Cited alongside, same era.
Hidden progress in deep learning: SGD learns parities near the computational limit
Boaz Barak, Benjamin Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang · 2022
Unifying grokking and double descent, 2023
Xander Davies, Lauro Langosco, and David Krueger · 2023
Closest in time.
Grokking as the transition from lazy to rich training dynamics, 2023
Tanishq Kumar, Blake Bordelon, Samuel J. Gershman, and Cengiz Pehlevan · 2023
Closest in time.
Grokking in linear estimators–a solvable model that groks without understanding
Noam Levi, Alon Beck, and Yohai Bar-Sinai · 2023
Closest in time.
Dichotomy of early and late phase implicit biases can provably induce grokking, 2023
Kaifeng Lyu, Jikai Jin, Zhiyuan Li, Simon S. Du, Jason D. Lee, and Wei Hu · 2023
Closest in time.
A tale of two circuits: Grokking as competition of sparse and dense subnetworks, 2023
William Merrill, Nikolaos Tsilivis, and Aman Shukla · 2023
Closest in time.
Grokking of hierarchical structure in vanilla transformers, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Towards understanding grokking: An effective theory of representation learning
Ziming Liu, Ouail Kitouni, Niklas S Nolte, Eric Michaud, Max Tegmark, and Mike Williams · 2022
Cited alongside, same era.
Grokking: Generalization beyond overfitting on small algorithmic datasets, 2022
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Cited alongside, same era.
Grokking phase transitions in learning local rules with gradient descent, 2022
Bojan Žunkovič and Enej Ilievski · 2022
Cited alongside, same era.
An Intuitive Explanation of Solomonoff Induction — LessWrong — lesswrong.com
Alex Altair · 2023
Cited alongside, same era.
Omnigrok: Grokking beyond algorithmic data
Ziming Liu, Eric J. Michaud, and Max Tegmark
Cited in the paper.
Grokking as compression: A nonlinear complexity perspective, 2023b
Ziming Liu, Ziqian Zhong, and Max Tegmark
Cited in the paper.
Shikhar Murty, Pratyusha Sharma, Jacob Andreas, and Christopher D. Manning · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability, 2023
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Closest in time.
Explaining grokking through circuit efficiency, 2023
Vikrant Varma, Rohin Shah, Zachary Kenton, János Kramár, and Ramana Kumar · 2023
Closest in time.
scipy.stats.pearsonr - scipy v1.11.4 manual
SciPy developers · 2024
Closest in time.