Fetching the paper…
Reading the bibliography…
We attribute grokking, the phenomenon where generalization is much delayed after memorization, to compression.
The information bottleneck method
Naftali Tishby, Fernando C Pereira, and William Bialek · 2000
Earlier work this paper cites.
A tutorial on spectral clustering
Ulrike Von Luxburg · 2007
Earlier work this paper cites.
Mathematische grundlagen der quantenmechanik , volume 38
John Von Neumann · 2013
Earlier work this paper cites.
On the number of linear regions of deep neural networks
Guido F Montufar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
On the expressive power of deep neural networks
Maithra Raghu, Ben Poole, Jon Kleinberg, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
Stefan Elfwing, Eiji Uchibe, and Kenji Doya · 2018
Earlier work this paper cites.
On the information bottleneck theory of deep learning
Andrew Michael Saxe, Yamini Bansal, Joel Dapello, Madhu Advani, Artemy Kolchinsky, Brendan Daniel Tracey, and David Daniel Cox · 2018
Earlier work this paper cites.
Complexity of linear regions in deep networks
Boris Hanin and David Rolnick · 2019
Cited alongside, same era.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2019
Cited alongside, same era.
Symbolic pregression: Discovering physical laws from distorted video
Silviu-Marian Udrescu and Max Tegmark · 2021
Cited alongside, same era.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Cited alongside, same era.
Hidden progress in deep learning: Sgd learns parities near the computational limit
Boaz Barak, Benjamin Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang · 2022
Cited alongside, same era.
A tale of two circuits: Grokking as competition of sparse and dense subnetworks
William Merrill, Nikolaos Tsilivis, and Aman Shukla · 2023
Closest in time.
Unifying grokking and double descent
Xander Davies, Lauro Langosco, and David Krueger · 2023
Closest in time.
Andrey Gromov · 2023
Closest in time.
Predicting grokking long before it happens: A look into the loss landscape of models which grok
Pascal Notsawo Jr, Hattie Zhou, Mohammad Pezeshki, Irina Rish, Guillaume Dumas, et al · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vimal Thilak, Etai Littwin, Shuangfei Zhai, Omid Saremi, Roni Paiss, and Joshua Susskind · 2022
Cited alongside, same era.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Liberum, Jess Smith, and Jacob Steinhardt · 2023
Cited alongside, same era.
Towards understanding grokking: An effective theory of representation learning
Ziming Liu, Ouail Kitouni, Niklas S Nolte, Eric Michaud, Max Tegmark, and Mike Williams
Cited in the paper.
Omnigrok: Grokking beyond algorithmic data
Ziming Liu, Eric J Michaud, and Max Tegmark
Cited in the paper.
Vikrant Varma, Rohin Shah, Zachary Kenton, János Kramár, and Ramana Kumar · 2023
Closest in time.
Language modeling is compression
Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, et al · 2023
Closest in time.