Fetching the paper…
Reading the bibliography…
We aim to understand grokking, a phenomenon where models generalize long after overfitting their training set.
Statistical mechanics for neural networks with continuous-time dynamics
R Kuhn and S Bos · 1993
Earlier work this paper cites.
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2013
Earlier work this paper cites.
Visual learning of arithmetic operation
Yedid Hoshen and Shmuel Peleg · 2016
Earlier work this paper cites.
Deep sets
Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Russ R Salakhutdinov, and Alexander J Smola · 2017
Earlier work this paper cites.
Training behavior of deep neural network in frequency domain
Zhi-Qin John Xu, Yaoyu Zhang, and Yanyang Xiao · 2019
Earlier work this paper cites.
Optimal regularization can mitigate double descent
Preetum Nakkiran, Prayaag Venkat, Sham Kakade, and Tengyu Ma · 2020
Earlier work this paper cites.
An overview of deep semi-supervised learning
Yassine Ouali, Céline Hudelot, and Myriam Tami · 2020
Earlier work this paper cites.
Bootstrap your own latent-a new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al · 2020
Earlier work this paper cites.
Contrastive representation learning: A framework and review
Phuc H Le-Khac, Graham Healy, and Alan F Smeaton · 2020
Earlier work this paper cites.
Neural mechanics: Symmetry and broken conservation laws in deep learning dynamics
Daniel Kunin, Javier Sagastuy-Brena, Surya Ganguli, Daniel LK Yamins, and Hidenori Tanaka · 2020
Earlier work this paper cites.
A free-energy principle for representation learning
Yansong Gao and Pratik Chaudhari · 2020
Earlier work this paper cites.
Generalisation error in learning with random features and the hidden manifold model
Federica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard, and Lenka Zdeborová · 2020
Earlier work this paper cites.
Prevalence of neural collapse during the terminal phase of deep learning training
Vardan Papyan, XY Han, and David L Donoho · 2020
Cited alongside, same era.
A type of generalization error induced by initialization in deep neural networks
Yaoyu Zhang, Zhi-Qin John Xu, Tao Luo, and Zheng Ma · 2020
Cited alongside, same era.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D. Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Cited alongside, same era.
Alignment Newsletter #159
Rohin Shah · 2021
Cited alongside, same era.
Machine-learning mathematical structures
Yang-Hui He · 2021
Cited alongside, same era.
Learning to unknot
Sergei Gukov, James Halverson, Fabian Ruehle, and Piotr Sułkowski · 2021
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Closest in time.
A mechanistic interpretability analysis of grokking, 2022
Neel Nanda and Tom Lieberum · 2022
Closest in time.
Grokking ’grokking’
Beren Millidge · 2022
Closest in time.
Multi-scale feature learning dynamics: Insights for double descent
Mohammad Pezeshki, Amartya Mitra, Yoshua Bengio, and Guillaume Lajoie · 2022
Closest in time.
The gaussian equivalence of generative models for learning with shallow neural networks
Sebastian Goldt, Bruno Loureiro, Galen Reeves, Florent Krzakala, Marc Mezard, and Lenka Zdeborova · 2022
Closest in time.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Advancing mathematics by guiding human intuition with ai
Alex Davies, Petar Veličković, Lars Buesing, Sam Blackwell, Daniel Zheng, Nenad Tomašev, Richard Tanburn, Peter Battaglia, Charles Blundell, András Juhász, et al · 2021
Cited alongside, same era.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2021
Cited alongside, same era.
Neural networks and quantum field theory
James Halverson, Anindita Maiti, and Keegan Stoner · 2021
Cited alongside, same era.
The principles of deep learning theory
Daniel A Roberts, Sho Yaida, and Boris Hanin · 2021
Cited alongside, same era.
Exploring simple siamese representation learning
Xinlei Chen and Kaiming He · 2021
Cited alongside, same era.
Closest in time.
The Principles of Deep Learning Theory
Daniel A. Roberts, Sho Yaida, and Boris Hanin · 2022
Closest in time.
Predictability and surprise in large generative models
Deep Ganguli, Danny Hernandez, Liane Lovitt, Nova DasSarma, Tom Henighan, Andy Jones, Nicholas Joseph, Jackson Kernion, Ben Mann, Amanda Askell, et al · 2022
Closest in time.
Future ML Systems Will Be Qualitatively Different
Jacob Steinhardt · 2022
Closest in time.
Thomson problem — Wikipedia, the free encyclopedia
Wikipedia contributors · 2022
Closest in time.
Omnigrok: Grokking beyond algorithmic data, 2022
Ziming Liu, Eric J. Michaud, and Max Tegmark · 2022
Closest in time.