Fetching the paper…
Reading the bibliography…
Grokking, the sudden generalization that occurs after prolonged overfitting, is a surprising phenomenon challenging our understanding of deep learning.
Incorporating second-order functional knowledge for better option pricing
Charles Dugas, Yoshua Bengio, François Bélisle, Claude Nadeau, and René Garcia · 2000
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research
Li Deng · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Pointnet: Deep learning on point sets for 3d classification and segmentation
Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas · 2017
Earlier work this paper cites.
On lazy training in differentiable programming
Lénaïc Chizat, Edouard Oyallon, and F. Bach · 2018
Earlier work this paper cites.
Ppfnet: Global context aware local features for robust 3d point matching
Haowen Deng, Tolga Birdal, and Slobodan Ilic · 2018
Earlier work this paper cites.
Risk and parameter convergence of logistic regression
Ziwei Ji and Matus Telgarsky · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Telgarsky · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Directional convergence and alignment in deep learning
Ziwei Ji and Matus Telgarsky · 2020
Earlier work this paper cites.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2020
Earlier work this paper cites.
Parameter norm growth during training of transformers
William Merrill, Vivek Ramanujan, Yoav Goldberg, Roy Schwartz, and Noah A. Smith · 2020
Cited alongside, same era.
Intrinsic dimension, persistent homology and generalization in neural networks
Tolga Birdal, Aaron Lou, Leonidas J Guibas, and Umut Simsekli · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Cited alongside, same era.
Adamp: Slowing down the slowdown for momentum optimizers on scale-invariant weights
Byeongho Heo, Sanghyuk Chun, Seong Joon Oh, Dongyoon Han, Sangdoo Yun, Gyuwan Kim, Youngjung Uh, and Jung-Woo Ha · 2021
Cited alongside, same era.
Hidden progress in deep learning: Sgd learns parities near the computational limit
Explaining grokking through circuit efficiency
Vikrant Varma, Rohin Shah, Zachary Kenton, János Kramár, and Ramana Kumar · 2023
Later among the works it cites.
Topological generalization bounds for discrete-time stochastic optimization algorithms
Rayna Andreeva, Benjamin Dupuis, Rik Sarkar, Tolga Birdal, and Umut Şimşekli · 2024
Later among the works it cites.
Grokking at the edge of linear separability
Alon Beck, Noam Levi, and Yohai Bar-Sinai · 2024
Later among the works it cites.
Deep networks always grok and here is why
Ahmed Imtiaz Humayun, Randall Balestriero, and Richard Baraniuk · 2024
Later among the works it cites.
Rotational equilibrium: How weight decay balances learning across neural networks, 2024
Atli Kosson, Bettina Messmer, and Martin Jaggi · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Boaz Barak, Benjamin Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang · 2022
Cited alongside, same era.
Deepstability: A study of unstable numerical methods and their solutions in deep learning
Eliska Kloberdanz, Kyle G Kloberdanz, and Wei Le · 2022
Cited alongside, same era.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Cited alongside, same era.
The slingshot mechanism: An empirical study of adaptive optimizers and the grokking phenomenon
Vimal Thilak, Etai Littwin, Shuangfei Zhai, Omid Saremi, Roni Paiss, and Josh Susskind · 2022
Cited alongside, same era.
A toy model of universality: Reverse engineering how networks learn group operations
Bilal Chughtai, Lawrence Chan, and Neel Nanda · 2023
Cited alongside, same era.
Why do we need weight decay in modern deep learning?
Francesco D’Angelo, Maksym Andriushchenko, Aditya Varre, and Nicolas Flammarion · 2023
Cited alongside, same era.
Andrey Gromov · 2023
Cited alongside, same era.
The asymmetric maximum margin bias of quasi-homogeneous neural networks
Daniel Kunin, Atsushi Yamamura, Chao Ma, and Surya Ganguli · 2023
Cited alongside, same era.
Later among the works it cites.
Grokking as the transition from lazy to rich training dynamics
Tanishq Kumar, Blake Bordelon, Samuel J. Gershman, and Cengiz Pehlevan · 2024
Later among the works it cites.
Language models” grok” to copy
Ang Lv, Ruobing Xie, Xingwu Sun, Zhanhui Kang, and Rui Yan · 2024
Later among the works it cites.
Dichotomy of early and late phase implicit biases can provably induce grokking
Kaifeng Lyu, Jikai Jin, Zhiyuan Li, Simon Shaolei Du, Jason D. Lee, and Wei Hu · 2024
Later among the works it cites.
Emergence in non-neural models: grokking modular arithmetic via average gradient outer product
Neil Mallinar, Daniel Beaglehole, Libin Zhu, Adityanarayanan Radhakrishnan, Parthe Pandit, and Mikhail Belkin · 2024
Later among the works it cites.
Grokking as a first order phase transition in two layer networks
Noa Rubin, Inbar Seroussi, and Zohar Ringel · 2024
Later among the works it cites.
Grokking group multiplication with cosets
Dashiell Stander, Qinan Yu, Honglu Fan, and Stella Biderman · 2024
Later among the works it cites.
Achieving margin maximization exponentially fast via progressive norm rescaling
Mingze Wang, Zeping Min, and Lei Wu · 2024
Later among the works it cites.
Grokking phase transitions in learning local rules with gradient descent
Bojan Žunkovič and Enej Ilievski · 2024
Later among the works it cites.