Fetching the paper…
Reading the bibliography…
Recently, an interesting phenomenon called grokking has gained much attention, where generalization occurs long after the models have initially overfitted the training data.
The information bottleneck method
Naftali Tishby, Fernando C Pereira, and William Bialek · 2000
Earlier work this paper cites.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky · 2015
Earlier work this paper cites.
On linear stability of sgd and input-smoothness of neural networks
Chao Ma and Lexing Ying · 2021
Earlier work this paper cites.
Information theory with kernel methods
Francis Bach · 2022
Earlier work this paper cites.
Hidden progress in deep learning: Sgd learns parities near the computational limit
Boaz Barak, Benjamin Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang · 2022
Earlier work this paper cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Earlier work this paper cites.
The slingshot mechanism: An empirical study of adaptive optimizers and the grokking phenomenon
Vimal Thilak, Etai Littwin, Shuangfei Zhai, Omid Saremi, Roni Paiss, and Joshua Susskind · 2022
Cited alongside, same era.
Grokking phase transitions in learning local rules with gradient descent
Bojan Žunkovič and Enej Ilievski · 2022
Cited alongside, same era.
Unifying grokking and double descent
Xander Davies, Lauro Langosco, and David Krueger · 2023
Cited alongside, same era.
Andrey Gromov · 2023
Cited alongside, same era.
A tale of two circuits: Grokking as competition of sparse and dense subnetworks
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Liberum, Jess Smith, and Jacob Steinhardt · 2023
Closest in time.
Predicting grokking long before it happens: A look into the loss landscape of models which grok
Pascal Notsawo Jr, Hattie Zhou, Mohammad Pezeshki, Irina Rish, Guillaume Dumas, et al · 2023
Closest in time.
Dime: Maximizing mutual information by a difference of matrix-based entropies
Oscar Skean, Jhoan Keider Hoyos Osorio, Austin J Brockmeier, and Luis Gonzalo Sanchez Giraldo · 2023
Closest in time.
Information flow in self-supervised learning
Zhiquan Tan, Jingqin Yang, Weiran Huang, Yang Yuan, and Yifan Zhang · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
William Merrill, Nikolaos Tsilivis, and Aman Shukla · 2023
Cited alongside, same era.
Towards understanding grokking: An effective theory of representation learning
Ziming Liu, Ouail Kitouni, Niklas S Nolte, Eric Michaud, Max Tegmark, and Mike Williams
Cited in the paper.
Omnigrok: Grokking beyond algorithmic data
Ziming Liu, Eric J Michaud, and Max Tegmark
Cited in the paper.
Matrix information theory for self-supervised learning
Yifan Zhang, Zhiquan Tan, Jingqin Yang, Weiran Huang, and Yang Yuan
Cited in the paper.
Relationmatch: Matching in-batch relationships for semi-supervised learning
Yifan Zhang, Jingqin Yang, Zhiquan Tan, and Yang Yuan
Cited in the paper.
Vikrant Varma, Rohin Shah, Zachary Kenton, János Kramár, and Ramana Kumar · 2023
Closest in time.