Fetching the paper…
Reading the bibliography…
Grokking, the unusual phenomenon for algorithmic datasets where generalization happens long after overfitting the training data, has remained elusive.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2011
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research
Li Deng · 2012
Earlier work this paper cites.
Quantum chemistry structures and properties of 134 kilo molecules
Raghunathan Ramakrishnan, Pavlo O Dral, Matthias Rupp, and O Anatole von Lilienfeld · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Samuel S Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein · 2016
Earlier work this paper cites.
Mean field residual networks: On the edge of chaos
Ge Yang and Samuel Schoenholz · 2017
Earlier work this paper cites.
Tunable efficient unitary neural networks (eunn) and their application to rnns
Li Jing, Yichen Shen, Tena Dubcek, John Peurifoy, Scott Skirlo, Yann LeCun, Max Tegmark, and Marin Soljačić · 2017
Earlier work this paper cites.
L2 regularization versus batch and weight normalization
Twan Van Laarhoven · 2017
Cited alongside, same era.
Three mechanisms of weight decay regularization
Guodong Zhang, Chaoqi Wang, Bowen Xu, and Roger Grosse · 2018
Cited alongside, same era.
The goldilocks zone: Towards better understanding of neural network loss landscapes
Stanislav Fort and Adam Scherlis · 2019
Cited alongside, same era.
Training behavior of deep neural network in frequency domain
Zhi-Qin John Xu, Yaoyu Zhang, and Yanyang Xiao · 2019
Cited alongside, same era.
Triple descent and the two kinds of overfitting: Where & why do they appear?
Stéphane d’Ascoli, Levent Sagun, and Giulio Biroli · 2020
Cited alongside, same era.
Alignment Newsletter #159
Rohin Shah · 2021
Later among the works it cites.
Multiple descent: Design your own generalization curve
Lin Chen, Yifei Min, Mikhail Belkin, and Amin Karbasi · 2021
Later among the works it cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Closest in time.
Towards understanding grokking: An effective theory of representation learning
Ziming Liu, Ouail Kitouni, Niklas Nolte, Eric J Michaud, Max Tegmark, and Mike Williams · 2022
Closest in time.
The slingshot mechanism: An empirical study of adaptive optimizers and the grokking phenomenon
Vimal Thilak, Etai Littwin, Shuangfei Zhai, Omid Saremi, Roni Paiss, and Joshua Susskind · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yasaman Bahri, Jonathan Kadmon, Jeffrey Pennington, Sam S Schoenholz, Jascha Sohl-Dickstein, and Surya Ganguli · 2020
Cited alongside, same era.
Revisiting initialization of neural networks
Maciej Skorski, Alessandro Temperoni, and Martin Theobald · 2020
Cited alongside, same era.
A type of generalization error induced by initialization in deep neural networks
Yaoyu Zhang, Zhi-Qin John Xu, Tao Luo, and Zheng Ma · 2020
Cited alongside, same era.
On the training dynamics of deep networks with l _ 2 l\_2 regularization
Aitor Lewkowycz and Guy Gur-Ari · 2020
Cited alongside, same era.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2021
Cited alongside, same era.
Hidden progress in deep learning: Sgd learns parities near the computational limit
Boaz Barak, Benjamin L Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang · 2022
Closest in time.
Cs229 lecture notes
Andrew Ng and Tengyu Ma · 2022
Closest in time.
Grokking ’grokking’
Beren Millidge · 2022
Closest in time.
Regularization-wise double descent: Why it occurs and how to eliminate it
Fatih Furkan Yilmaz and Reinhard Heckel · 2022
Closest in time.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Liberum, Jess Smith, and Jacob Steinhardt · 2023
Closest in time.