Fetching the paper…
Reading the bibliography…
Grokking is a phenomenon where a model trained on an algorithmic task first overfits but, then, after a large amount of additional training, undergoes a phase transition to generalize perfectly.
Skeletonization: A Technique for Trimming the Fat from a Network via Relevance Assessment , pp. 107–115
Michael C. Mozer and Paul Smolensky · 1989
Earlier work this paper cites.
Statistical Mechanics of Learning
A. Engel and C. Van den Broeck · 2001
Earlier work this paper cites.
Effects of parameter norm growth during transformer training: Inductive bias from gradient descent
William Merrill, Vivek Ramanujan, Yoav Goldberg, Roy Schwartz, and Noah A. Smith · 2021
Earlier work this paper cites.
Hidden progress in deep learning: SGD learns parities near the computational limit
Boaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade, Eran Malach, and Cyril Zhang · 2022
Earlier work this paper cites.
GPT3.int8(): 8-bit matrix multiplication for transformers at scale
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer · 2022
Cited alongside, same era.
Grokking: Generalization beyond overfitting on small algorithmic datasets, 2022
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Cited alongside, same era.
The slingshot mechanism: An empirical study of adaptive optimizers and the \emph{Grokking Phenomenon}
Vimal Thilak, Etai Littwin, Shuangfei Zhai, Omid Saremi, Roni Paiss, and Joshua M. Susskind · 2022
Cited alongside, same era.
Emergent abilities of large language models
Barret Zoph, Colin Raffel, Dale Schuurmans, Dani Yogatama, Denny Zhou, Don Metzler, Ed H. Chi, Jason Wei, Jeff Dean, Liam B. Fedus, Maarten Paul Bosma, Oriol Vinyals, Percy Liang, Sebastian Borgeaud, Tatsunori B. Hashimoto, and Yi Tay · 2022
Later among the works it cites.
Omnigrok: Grokking beyond algorithmic data
Ziming Liu, Eric J Michaud, and Max Tegmark · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…