Fetching the paper…
Reading the bibliography…
We present a theoretical explanation of the ``grokking'' phenomenon, where a model generalizes long after overfitting,for the originally-studied problem of modular addition.
“Wide neural networks of any depth evolve as linear models under gradient descent”
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein and Jeffrey Pennington · 1902
Earlier work this paper cites.
“On exact computation with an infinitely wide neural net”
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, Russ Salakhutdinov and Ruosong Wang · 1904
Earlier work this paper cites.
“Implicit regularization for deep neural networks driven by an Ornstein-Uhlenbeck like process”
Guy Blanc, Neha Gupta, Gregory Valiant and Paul Valiant · 1904
Earlier work this paper cites.
“Implicit Regularization in Deep Matrix Factorization”
Sanjeev Arora, Nadav Cohen, Wei Hu and Yuping Luo · 1905
Earlier work this paper cites.
“Lexicographic and Depth-Sensitive Margins in Homogeneous and Non-Homogeneous Deep Models”
Mor Nacson, Suriya Gunasekar, Jason. Lee, Nathan Srebro and Daniel Soudry · 1905
Earlier work this paper cites.
“Disentangling feature and lazy training in deep neural networks”
Mario Geiger, Stefano Spigler, Arthur Jacot and Matthieu Wyart · 1906
Earlier work this paper cites.
“Gradient descent maximizes the margin of homogeneous neural networks”
Kaifeng Lyu and Jian Li · 1906
Earlier work this paper cites.
“Methods of Mathematical Physics”
Richard Courant and David Hilbert · 1953
Earlier work this paper cites.
“‘Balls into Bins’ — A Simple and Tight Analysis”
Martin Raab and Angelika Steger · 1998
Earlier work this paper cites.
“Simplified PAC-Bayesian Margin Bounds”
David. McAllester · 2003
Earlier work this paper cites.
“Feature selection, L 1 L_{1} vs. L 2 L_{2} regularization, and rotational invariance”
Andrew Ng · 2004
Earlier work this paper cites.
“Shape Matters: Understanding the Implicit Bias of the Noise Covariance”
Jeff. HaoChen, Colin Wei, Jason. Lee and Tengyu Ma · 2006
Earlier work this paper cites.
“Implicit Bias in Deep Linear Classification: Initialization Scale vs Training Accuracy”
Edward Moroshko, Suriya Gunasekar, Blake Woodworth, Jason. Lee, Nathan Srebro and Daniel Soudry · 2007
Earlier work this paper cites.
“When Hardness of Approximation Meets Hardness of Learning”
Eran Malach and Shai Shalev-Shwartz · 2008
Earlier work this paper cites.
Stanislav Fort, Gintare Dziugaite, Mansheej Paul, Sepideh Kharaghani, Daniel. Roy and Surya Ganguli · 2010
Earlier work this paper cites.
“Why are convolutional nets more sample-efficient than fully-connected nets?”
Zhiyuan Li, Yi Zhang and Sanjeev Arora · 2010
Earlier work this paper cites.
“Smoothness, low noise and fast rates”
Nathan Srebro, Karthik Sridharan and Ambuj Tewari · 2010
Earlier work this paper cites.
“Tensor Programs IV: Feature Learning in Infinite-Width Neural Networks”
Greg Yang and Edward. Hu · 2011
Earlier work this paper cites.
“Kernels for vector-valued functions: A review”
Mauricio Álvarez, Lorenzo Rosasco and Neil Lawrence · 2012
Earlier work this paper cites.
“Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2015
Cited alongside, same era.
“An Introduction to Matrix Concentration Inequalities”
Joel. Tropp · 2015
Cited alongside, same era.
“Implicit Regularization in Matrix Factorization”
Suriya Gunasekar, Blake Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur and Nathan Srebro · 2017
Cited alongside, same era.
“JAX: composable transformations of Python+NumPy programs”, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne and Qiao Zhang · 2018
Cited alongside, same era.
“Characterizing Implicit Bias in Terms of Optimization Geometry”
Vimal Thilak, Etai Littwin, Shuangfei Zhai, Omid Saremi, Roni Paiss and Joshua Susskind · 2022
Later among the works it cites.
“Simplicity Bias in Transformers and their Ability to Learn Sparse Boolean Functions”
Satwik Bhattamishra, Arkil Patel, Varun Kanade and Phil Blunsom · 2023
Later among the works it cites.
“Grokking modular arithmetic”, 2023
Andrey Gromov · 2023
Later among the works it cites.
“Omnigrok: Grokking beyond algorithmic data”
Ziming Liu, Eric Michaud and Max Tegmark · 2023
Later among the works it cites.
“A fast, well-founded approximation to the empirical neural tangent kernel”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Suriya Gunasekar, Jason Lee, Daniel Soudry and Nathan Srebro · 2018
Cited alongside, same era.
“Neural tangent kernel: Convergence and generalization in neural networks”
Arthur Jacot, Franck Gabriel and Cl\’ement Hongler · 2018
Cited alongside, same era.
“Foundations of Machine Learning”
Mehryar Mohri, Afshin Rostamizadeh and Ameet Talkwalkar · 2018
Cited alongside, same era.
“A PAC-Bayesian Approach to Spectrally-Normalized Margin Bounds for Neural Networks”
Behnam Neyshabur, Srinadh Bhojanapalli and Nathan Srebro · 2018
Cited alongside, same era.
“The Implicit Bias of Gradient Descent on Separable Data”
Daniel Soudry, Elad Hoffer, Mor Nacson, Suriya Gunasekar and Nathan Srebro · 2018
Cited alongside, same era.
“On Lazy Training in Differentiable Programming”
Lenaic Chizat, Edouard Oyallon and Francis Bach · 2019
Cited alongside, same era.
“High-Dimensional Statistics: A Non-Asymptotic Viewpoint”
Martin. Wainwright · 2019
Cited alongside, same era.
“Regularization Matters: Generalization and Optimization of Neural Nets v.s. their Induced Kernel”
Colin Wei, Jason. Lee, Qiang Liu and Tengyu Ma · 2019
Cited alongside, same era.
Mohamad Mohamadi, Wonho Bae and Danica Sutherland · 2023
Later among the works it cites.
“Progress measures for grokking via mechanistic interpretability”
Neel Nanda, Lawrence Chan, Tom Liberum, Jess Smith and Jacob Steinhardt · 2023
Later among the works it cites.
Pascal Jr. Notsawo, Hattie Zhou, Mohammad Pezeshki, Irina Rish and Guillaume Dumas · 2023
Later among the works it cites.
“Explaining grokking through circuit efficiency”, 2023
Vikrant Varma, Rohin Shah, Zachary Kenton, János Kramár and Ramana Kumar · 2023
Later among the works it cites.
“PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation”
Jason Ansel, Edward Yang, Horace He, Natalia Gimelshein, Animesh Jain, Michael Voznesensky, Bin Bao, Peter Bell, David Berard, Evgeni Burovski, Geeta Chauhan, Anjali Chourdia, Will Constable, Alban Desmaison, Zachary DeVito, Elias Ellison, Will Feng, Jiong Gong, Michael Gschwind, Brian Hirsh, Sherlock Huang, Kshiteej Kalambarkar, Laurent Kirsch, Michael Lazos, Mario Lezcano, Yanbo Liang, Jason Liang, Yinghai Lu, C.. Luk, Bert Maher, Yunjie Pan, Christian Puhrsch, Matthias Reso, Mark Saroufim, Marcos Siraichi, Helen Suk, Shunting Zhang, Michael Suo, Phil Tillet, Xu Zhao, Eikan Wang, Keren Zhou, Richard Zou, Xiaodong Wang, Ajit Mathews, William Wen, Gregory Chanan, Peng Wu and Soumith Chintala · 2024
Closest in time.
“Learning the greatest common divisor: explaining transformer predictions”
François Charton · 2024
Closest in time.
“Grokking as the Transition from Lazy to Rich Training Dynamics”
Tanishq Kumar, Blake Bordelon, Samuel. Gershman and Cengiz Pehlevan · 2024
Closest in time.
“Grokking in Linear Estimators – A Solvable Model that Groks without Understanding”
Noam Levi, Alon Beck and Yohai Bar-Sinai · 2024
Closest in time.
“Dichotomy of Early and Late Phase Implicit Biases Can Provably Induce Grokking”
Kaifeng Lyu, Jikai Jin, Zhiyuan Li, Simon. Du, Jason. Lee and Wei Hu · 2024
Closest in time.
“Feature emergence via margin maximization: case studies in algebraic tasks”
Depen Morwani, Benjamin. Edelman, Costin-Andrei Oncescu, Rosie Zhao and Sham Kakade · 2024
Closest in time.
“Grokking as a First Order Phase Transition in Two Layer Networks”
Noa Rubin, Inbar Seroussi and Zohar Ringel · 2024
Closest in time.
“Implicit Bias of AdamW: ℓ ∞ \ell_{\infty} Norm Constrained Optimization”
Shuo Xie and Zhiyuan Li · 2024
Closest in time.
“Benign Overfitting and Grokking in ReLU Networks for XOR Cluster Data”
Zhiwei Xu, Yutong Wang, Spencer Frei, Gal Vardi and Wei Hu · 2024
Closest in time.
“The Implicit Bias of Adam on Separable Data”, 2024
Chenyang Zhang, Difan Zou and Yuan Cao · 2024
Closest in time.