Fetching the paper…
Reading the bibliography…
Grokking is the intriguing phenomenon where a model learns to generalize long after it has fit the training data.
Distribution of eigenvalues for some sets of random matrices
V A Marčenko and L A Pastur · 1967
Earlier work this paper cites.
A simple weight decay can improve generalization
Anders Krogh and John A. Hertz · 1991
Earlier work this paper cites.
Statistical mechanics of learning from examples
Hyunjune Sebastian Seung, Haim Sompolinsky, and Naftali Tishby · 1992
Earlier work this paper cites.
High-dimensional asymptotics of prediction: Ridge regression and classification, 2015
Edgar Dobriban and Stefan Wager · 2015
Earlier work this paper cites.
Dynamics of learning in deep linear neural networks: A mean-field approach
Andrea Crisanti and Haim Sompolinsky · 2018
Earlier work this paper cites.
Three mechanisms of weight decay regularization
Guodong Zhang, Chaoqi Wang, Bowen Xu, and Roger B. Grosse · 2018
Earlier work this paper cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Earlier work this paper cites.
Modelling the infinite width limit of neural networks with mean field theory
Sebastian Goldt, Marc M’ezard, Florent Krzakala, and Lenka Zdeborov’a · 2020
Earlier work this paper cites.
Dynamics of generalization in learning with gradient descent for piecewise linear neural networks
Eric Bodin and Nicolas Macris · 2021
Cited alongside, same era.
Asymptotics of ridge (less) regression under general source condition, 2021
Dominic Richards, Jaouad Mourtada, and Lorenzo Rosasco · 2021
Cited alongside, same era.
Gradient flow in the gaussian covariate model: exact solution of learning curves and multiple descent structures, 2022
Antoine Bodin and Nicolas Macris · 2022
Cited alongside, same era.
Towards understanding grokking: An effective theory of representation learning
Ziming Liu, Ouail Kitouni, Niklas S Nolte, Eric Michaud, Max Tegmark, and Mike Williams · 2022
Cited alongside, same era.
Learning curves of generic features maps for realistic datasets with a teacher-student model
Bruno Loureiro, Cedric Gerbelot, Hugo Cui, Sebastian Goldt, Florent Krzakala, Marc Mezard, and Lenka Zdeborova · 2022
Cited alongside, same era.
Grokking phase transitions in learning local rules with gradient descent, 2022
Bojan Žunkovič and Enej Ilievski · 2022
Later among the works it cites.
A toy model of universality: Reverse engineering how networks learn group operations, 2023
Bilal Chughtai, Lawrence Chan, and Neel Nanda · 2023
Closest in time.
Unifying grokking and double descent
Xander Davies, Lauro Langosco, and David Krueger · 2023
Closest in time.
Omnigrok: Grokking beyond algorithmic data, 2023
Ziming Liu, Eric J. Michaud, and Max Tegmark · 2023
Closest in time.
A tale of two circuits: Grokking as competition of sparse and dense subnetworks, 2023
William Merrill, Nikolaos Tsilivis, and Aman Shukla · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Grokking ’grokking’, 2022
Beren Millidge · 2022
Cited alongside, same era.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Cited alongside, same era.
The slingshot mechanism: An empirical study of adaptive optimizers and the grokking phenomenon
Vimal Thilak, Etai Littwin, Shuangfei Zhai, Omid Saremi, Roni Paiss, and Joshua Susskind · 2022
Cited alongside, same era.
Neel Nanda, Lawrence Chan, Tom Liberum, Jess Smith, and Jacob Steinhardt · 2023
Closest in time.
Predicting grokking long before it happens: A look into the loss landscape of models which grok, 2023
Pascal Jr. Tikeng Notsawo, Hattie Zhou, Mohammad Pezeshki, Irina Rish, and Guillaume Dumas · 2023
Closest in time.