Fetching the paper…
Reading the bibliography…
Neural networks readily learn a subset of the modular arithmetic tasks, while failing to generalize on the rest.
Hierarchical mixtures of experts and the em algorithm
M.I. Jordan and R.A. Jacobs · 1993
Earlier work this paper cites.
On lattices, learning with errors, random linear codes, and cryptography
Oded Regev · 2009
Earlier work this paper cites.
Modular arithmetic and microtonal music theory
K. Kennedy A. AsKew and V. Klima · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, *Azalia Mirhoseini, *Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
High performance simd modular arithmetic for polynomial evaluation, 2020
Pierre Fortin, Ambroise Fleury, François Lemaire, and Michael Monagan · 2020
Earlier work this paper cites.
{GS}hard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2021
Earlier work this paper cites.
Hidden progress in deep learning: SGD learns parities near the computational limit
Boaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade, eran malach, and Cyril Zhang · 2022
Earlier work this paper cites.
Unifying grokking and double descent
Xander Davies, Lauro Langosco, and David Krueger · 2022
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2022
William Fedus, Barret Zoph, and Noam Shazeer · 2022
Earlier work this paper cites.
Towards understanding grokking: An effective theory of representation learning
Ziming Liu, Ouail Kitouni, Niklas S Nolte, Eric Michaud, Max Tegmark, and Mike Williams · 2022
Cited alongside, same era.
Grokking: Generalization beyond overfitting on small algorithmic datasets, 2022
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Cited alongside, same era.
The slingshot mechanism: An empirical study of adaptive optimizers and the grokking phenomenon, 2022
Vimal Thilak, Etai Littwin, Shuangfei Zhai, Omid Saremi, Roni Paiss, and Joshua Susskind · 2022
Cited alongside, same era.
To grok or not to grok: Disentangling generalization and memorization on corrupted algorithmic datasets, 2023
Darshil Doshi, Aritra Das, Tianyu He, and Andrey Gromov · 2023
Cited alongside, same era.
Grokking modular arithmetic, 2023
Andrey Gromov · 2023
Cited alongside, same era.
A tale of two circuits: Grokking as competition of sparse and dense subnetworks, 2023
William Merrill, Nikolaos Tsilivis, and Aman Shukla · 2023
Later among the works it cites.
Grokking tickets: Lottery tickets accelerate grokking, 2023
Gouki Minegishi, Yusuke Iwasawa, and Yutaka Matsuo · 2023
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability, 2023
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Later among the works it cites.
Predicting grokking long before it happens: A look into the loss landscape of models which grok, 2023
Pascal Jr. Tikeng Notsawo, Hattie Zhou, Mohammad Pezeshki, Irina Rish, and Guillaume Dumas · 2023
Later among the works it cites.
Droplets of good representations: Grokking as a first order phase transition in two layer networks, 2023
Noa Rubin, Inbar Seroussi, and Zohar Ringel · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Grokking as the transition from lazy to rich training dynamics, 2023
Tanishq Kumar, Blake Bordelon, Samuel J. Gershman, and Cengiz Pehlevan · 2023
Cited alongside, same era.
Omnigrok: Grokking beyond algorithmic data
Ziming Liu, Eric J Michaud, and Max Tegmark · 2023
Cited alongside, same era.
Dichotomy of early and late phase implicit biases can provably induce grokking, 2023
Kaifeng Lyu, Jikai Jin, Zhiyuan Li, Simon S. Du, Jason D. Lee, and Wei Hu · 2023
Cited alongside, same era.
Explaining grokking through circuit efficiency, 2023
Vikrant Varma, Rohin Shah, Zachary Kenton, János Kramár, and Ramana Kumar · 2023
Later among the works it cites.
Salsa: Attacking lattice cryptography with transformers, 2023
Emily Wenger, Mingjie Chen, François Charton, and Kristin Lauter · 2023
Later among the works it cites.
The clock and the pizza: Two stories in mechanistic explanation of neural networks, 2023
Ziqian Zhong, Ziming Liu, Max Tegmark, and Jacob Andreas · 2023
Later among the works it cites.