Fetching the paper…
Reading the bibliography…
We investigate the phenomenon of grokking -- delayed generalization accompanied by non-monotonic test loss behavior -- in a simple binary logistic classification task, for which "memorizing" and "generalizing" solutions can be strictly defined.
A problem in geometric probability
Wendel, J · 1962
Earlier work this paper cites.
Grokking phase transitions in learning local rules with gradient descent, 2022
Žunkovič, B. and Ilievski, E · 1962
Earlier work this paper cites.
Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition
Cover, T. M · 1965
Earlier work this paper cites.
Inflationary universe: A possible solution to the horizon and flatness problems
Guth, A. H · 1981
Earlier work this paper cites.
Early stopping-but when?
Prechelt, L · 1996
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
Kingma, D. P. and Ba, J · 2017
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on nonseparable data
Ji, Z. and Telgarsky, M · 2019
Earlier work this paper cites.
Convergence of gradient descent on separable data, 2019
Nacson, M. S., Lee, J. D., Gunasekar, S., Savarese, P. H. P., Srebro, N., and Soudry, D · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
The large learning rate phase of deep learning: the catapult mechanism, 2020
Lewkowycz, A., Bahri, Y., Dyer, E., Sohl-Dickstein, J., and Gur-Ari, G · 2020
Earlier work this paper cites.
Statistical mechanics: entropy, order parameters, and complexity , volume 14
Sethna, J. P · 2021
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Earlier work this paper cites.
Gradient descent on neural networks typically occurs at the edge of stability, 2022
Cohen, J. M., Kaur, S., Li, Y., Kolter, J. Z., and Talwalkar, A · 2022
Cited alongside, same era.
Towards understanding grokking: An effective theory of representation learning
Liu, Z., Kitouni, O., Nolte, N. S., Michaud, E., Tegmark, M., and Williams, M · 2022
Cited alongside, same era.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Power, A., Burda, Y., Edwards, H., Babuschkin, I., and Misra, V · 2022
Cited alongside, same era.
The slingshot mechanism: An empirical study of adaptive optimizers and the grokking phenomenon
Thilak, V., Littwin, E., Zhai, S., Saremi, O., Paiss, R., and Susskind, J · 2022
Cited alongside, same era.
Progress measures for grokking via mechanistic interpretability
Nanda, N., Chan, L., Liberum, T., Smith, J., and Steinhardt, J · 2023
Later among the works it cites.
Predicting grokking long before it happens: A look into the loss landscape of models which grok, 2023
Notsawo, P. J. T., Zhou, H., Pezeshki, M., Rish, I., and Dumas, G · 2023
Later among the works it cites.
Droplets of good representations: Grokking as a first order phase transition in two layer networks, 2023
Rubin, N., Seroussi, I., and Ringel, Z · 2023
Later among the works it cites.
Double descent demystified: Identifying, interpreting & ablating the sources of a deep learning puzzle, 2023
Schaeffer, R., Khona, M., Robertson, Z., Boopathy, A., Pistunova, K., Rocks, J. W., Fiete, I. R., and Koyejo, O · 2023
Later among the works it cites.
Omnivec: Learning robust representations with cross modal sharing, 2023
Srivastava, S. and Sharma, G · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zeng, A., Liu, X., Du, Z., Wang, Z., Lai, H., Ding, M., Yang, Z., Xu, Y., Zheng, W., Xia, X., et al · 2022
Cited alongside, same era.
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al · 2023
Cited alongside, same era.
A toy model of universality: Reverse engineering how networks learn group operations, 2023
Chughtai, B., Chan, L., and Nanda, N · 2023
Cited alongside, same era.
Unifying grokking and double descent
Davies, X., Langosco, L., and Krueger, D · 2023
Cited alongside, same era.
Bard - chat based ai tool from google (october 2023 version) [large language model]
Google · 2023
Cited alongside, same era.
Gromov, A · 2023
Cited alongside, same era.
Grokking as the transition from lazy to rich training dynamics, 2023
Kumar, T., Bordelon, B., Gershman, S. J., and Pehlevan, C · 2023
Cited alongside, same era.
Grokking in linear estimators – a solvable model that groks without understanding, 2023
Levi, N., Beck, A., and Bar-Sinai, Y · 2023
Cited alongside, same era.
Attention is all you need, 2023
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2023
Later among the works it cites.
Benign overfitting and grokking in relu networks for xor cluster data, 2023
Xu, Z., Wang, Y., Frei, S., Vardi, G., and Hu, W · 2023
Later among the works it cites.
To grok or not to grok: Disentangling generalization and memorization on corrupted algorithmic datasets, 2024
Doshi, D., Das, A., He, T., and Gromov, A · 2024
Closest in time.
Progress measures for grokking on real-world tasks, 2024
Golechha, S · 2024
Closest in time.
Deep networks always grok and here is why, 2024
Humayun, A. I., Balestriero, R., and Baraniuk, R · 2024
Closest in time.
Dichotomy of early and late phase implicit biases can provably induce grokking, 2024
Lyu, K., Jin, J., Li, Z., Du, S. S., Lee, J. D., and Hu, W · 2024
Closest in time.
OpenAI · 2024
Closest in time.
Grokking as a first order phase transition in two layer networks, 2024
Rubin, N., Seroussi, I., and Ringel, Z · 2024
Closest in time.
Grokking at the edge of numerical stability, 2025
Prieto, L., Barsbey, M., Mediano, P. A. M., and Birdal, T · 2025
Closest in time.