Fetching the paper…
Reading the bibliography…
Neural networks trained to solve modular arithmetic tasks exhibit grokking, a phenomenon where the test accuracy starts improving long after the model achieves 100% training accuracy in the training process.
A course in number theory and cryptography
N. Koblitz · 1994
Earlier work this paper cites.
Structure adaptive approach for dimension reduction
M. Hristache, A. Juditsky, J. Polzehl, and V. Spokoiny · 2001
Earlier work this paper cites.
Toeplitz and circulant matrices: A review
R. M. Gray et al · 2006
Earlier work this paper cites.
A consistent estimator of the expected gradient outerproduct
S. Trivedi, J. Wang, S. Kpotufe, and G. Shakhnarovich · 2014
Earlier work this paper cites.
Implicit regularization in matrix factorization
S. Gunasekar, B. E. Woodworth, S. Bhojanapalli, B. Neyshabur, and N. Srebro · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Foundations of machine learning
M. Mohri, A. Rostamizadeh, and A. Talwalkar · 2018
Earlier work this paper cites.
Algorithmic aspects of machine learning
A. Moitra · 2018
Earlier work this paper cites.
Robust learning with jacobian regularization
J. Hoffman, D. A. Roberts, and S. Yaida · 2019
Earlier work this paper cites.
Deep learning: a statistical viewpoint
P. L. Bartlett, A. Montanari, and A. Rakhlin · 2021
Earlier work this paper cites.
Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation
M. Belkin · 2021
Earlier work this paper cites.
Hidden progress in deep learning: Sgd learns parities near the computational limit
B. Barak, B. Edelman, S. Goel, S. Kakade, E. Malach, and C. Zhang · 2022
Earlier work this paper cites.
Neural networks can learn representations with gradient descent
A. Damian, J. Lee, and M. Soltanolkotabi · 2022
Earlier work this paper cites.
Towards understanding grokking: An effective theory of representation learning
Z. Liu, O. Kitouni, N. S. Nolte, E. Michaud, M. Tegmark, and M. Williams · 2022
Earlier work this paper cites.
Neural networks efficiently learn low-dimensional representations with sgd
A. Mousavi-Hosseini, S. Park, M. Girotti, I. Mitliagkas, and M. A. Erdogdu · 2022
Earlier work this paper cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets
A. Power, Y. Burda, H. Edwards, I. Babuschkin, and V. Misra · 2022
Earlier work this paper cites.
A. Radhakrishnan, D. Beaglehole, P. Pandit, and M. Belkin · 2022
Cited alongside, same era.
The slingshot mechanism: An empirical study of adaptive optimizers and the grokking phenomenon
V. Thilak, E. Littwin, S. Zhai, O. Saremi, R. Paiss, and J. Susskind · 2022
Cited alongside, same era.
Understanding the covariance structure of convolutional filters
A. Trockman, D. Willmott, and J. Z. Kolter · 2022
Cited alongside, same era.
Emergent abilities of large language models
J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus · 2022
Cited alongside, same era.
Intrinsic dimensionality and generalization properties of the r-norm inductive bias
Are emergent abilities of large language models a mirage?
R. Schaeffer, B. Miranda, and S. Koyejo · 2023
Later among the works it cites.
Explaining grokking through circuit efficiency
V. Varma, R. Shah, Z. Kenton, J. Kramár, and R. Kumar · 2023
Later among the works it cites.
Efficient estimation of the central mean subspace via smoothed gradient outer products
G. Yuan, M. Xu, S. Kpotufe, and D. Hsu · 2023
Later among the works it cites.
D. Beaglehole, I. Mitliagkas, and A. Agarwala · 2024
Closest in time.
Average gradient outer product as a mechanism for deep neural collapse
D. Beaglehole, P. Súkeník, M. Mondelli, and M. Belkin · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Ardeshir, D. J. Hsu, and C. H. Sanford · 2023
Cited alongside, same era.
A theory for emergence of complex skills in language models
S. Arora and A. Goyal · 2023
Cited alongside, same era.
Mechanism of feature learning in convolutional neural networks
D. Beaglehole, A. Radhakrishnan, P. Pandit, and M. Belkin · 2023
Cited alongside, same era.
Unifying grokking and double descent
X. Davies, L. Langosco, and D. Krueger · 2023
Cited alongside, same era.
Agnostically learning multi-index models with queries
I. Diakonikolas, D. M. Kane, V. Kontonis, C. Tzamos, and N. Zarifis · 2023
Cited alongside, same era.
A. Gromov · 2023
Cited alongside, same era.
Omnigrok: Grokking beyond algorithmic data
Z. Liu, E. J. Michaud, and M. Tegmark · 2023
Cited alongside, same era.
Dichotomy of early and late phase implicit biases can provably induce grokking
K. Lyu, J. Jin, Z. Li, S. S. Du, J. D. Lee, and W. Hu · 2023
Cited alongside, same era.
Grokking modular polynomials
D. Doshi, T. He, A. Das, and A. Gromov · 2024
Closest in time.
Interpreting grokked transformers in complex modular arithmetic
H. Furuta, G. Minegishi, Y. Iwasawa, and Y. Matsuo · 2024
Closest in time.
Grokking as the transition from lazy to rich training dynamics
T. Kumar, B. Bordelon, S. J. Gershman, and C. Pehlevan · 2024
Closest in time.
Grokking beyond neural networks: An empirical exploration with model complexity
J. Miller, C. O’Neill, and T. Bui · 2024
Closest in time.
Why do you grok? a theoretical analysis on grokking modular addition
M. A. Mohamadi, Z. Li, L. Wu, and D. J. Sutherland · 2024
Closest in time.
Feature emergence via margin maximization: case studies in algebraic tasks
D. Morwani, B. L. Edelman, C.-A. Oncescu, R. Zhao, and S. Kakade · 2024
Closest in time.
Mechanism for feature learning in neural networks and backpropagation-free machine learning models
A. Radhakrishnan, D. Beaglehole, P. Pandit, and M. Belkin · 2024
Closest in time.
Linear recursive feature machines provably recover low-rank matrices
A. Radhakrishnan, M. Belkin, and D. Drusvyatskiy · 2024
Closest in time.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
Z. Zhong, Z. Liu, M. Tegmark, and J. Andreas · 2024
Closest in time.
Catapults in sgd: spikes in the training loss and their impact on generalization through feature learning
L. Zhu, C. Liu, A. Radhakrishnan, and M. Belkin · 2024
Closest in time.