Fetching the paper…
Reading the bibliography…
This paper focuses on predicting the occurrence of grokking in neural networks, a phenomenon in which perfect generalization emerges long after signs of overfitting or memorization are observed.
Some methods of speeding up the convergence of iteration methods
Polyak, B · 1964
Earlier work this paper cites.
Eeg analysis based on time domain properties
Hjorth, B · 1970
Earlier work this paper cites.
A direct adaptive method for faster backpropagation learning: the rprop algorithm
Riedmiller, M. and Braun, H · 1993
Earlier work this paper cites.
Flat Minima
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Nonlinear dimensionality reduction by locally linear embedding
Roweis, S. T. and Saul, L. K · 2000
Earlier work this paper cites.
A global geometric framework for nonlinear dimensionality reduction
Tenenbaum, J. B., de Silva, V., and Langford, J. C · 2000
Earlier work this paper cites.
Laplacian eigenmaps and spectral techniques for embedding and clustering
Belkin, M. and Niyogi, P · 2001
Earlier work this paper cites.
Hessian eigenmaps: Locally linear embedding techniques for high-dimensional data
Donoho, D. L. and Grimes, C · 2003
Earlier work this paper cites.
Maximum likelihood estimation of intrinsic dimension
Levina, E. and Bickel, P · 2004
Earlier work this paper cites.
Comments on ‘maximum likelihood estimation of intrinsic dimension’ by elizaveta levina and peter bickel (2004)
David J.C., M. and Zoubin, G · 2005
Earlier work this paper cites.
Locating and characterizing the stationary points of the extended rosenbrock function
Kok, S. and Sandrock, C · 2009
Earlier work this paper cites.
Neural networks for machine learning. coursera, video lectures, 2012
Hinton, G · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Graves, A., Mohamed, A., and Hinton, G. E · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
Livni, R., Shalev-Shwartz, S., and Shamir, O · 2014
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Bubeck, S. et al · 2015
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Goodfellow, I. J. and Vinyals, O · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Cited alongside, same era.
An empirical analysis of the optimization of deep network loss surfaces
Im, D. J., Tao, M., and Branson, K · 2016
Cited alongside, same era.
Gradient descent converges to minimizers
Lee, J. D., Simchowitz, M., Jordan, M. I., and Recht, B · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D · 2016
Cited alongside, same era.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y · 2017
Cited alongside, same era.
Fantastic generalization measures and where to find them
Jiang, Y., Neyshabur, B., Mobahi, H., Krishnan, D., and Bengio, S · 2019
Later among the works it cites.
The large learning rate phase of deep learning: the catapult mechanism
Lewkowycz, A., Bahri, Y., Dyer, E., Sohl-Dickstein, J., and Gur-Ari, G · 2020
Later among the works it cites.
Understanding the role of training regimes in continual learning
Mirzadeh, S. I., Farajtabar, M., Pascanu, R., and Ghasemzadeh, H · 2020
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I · 2020
Later among the works it cites.
The modern mathematics of deep learning
Berner, J., Grohs, P., Kutyniok, G., and Petersen, P · 2021
Later among the works it cites.
Multiple descent: Design your own generalization curve
Chen, L., Min, Y., Belkin, M., and Karbasi, A · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
Dziugaite, G. K. and Roy, D. M · 2017
Cited alongside, same era.
Estimating the intrinsic dimension of datasets by a minimal neighborhood information
Facco, E., d’Errico, M., Rodriguez, A., and Laio, A · 2017
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2017
Cited alongside, same era.
Exploring generalization in deep learning
Neyshabur, B., Bhojanapalli, S., McAllester, D., and Srebro, N · 2017
Cited alongside, same era.
Opening the black box of deep neural networks via information
Shwartz-Ziv, R. and Tishby, N · 2017
Cited alongside, same era.
Exploring loss function topology with cyclical learning rates
Smith, L. N. and Topin, N · 2017
Cited alongside, same era.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
Later among the works it cites.
Gradient descent on neural networks typically occurs at the edge of stability
Cohen, J. M., Kaur, S., Li, Y., Kolter, J. Z., and Talwalkar, A. S · 2021
Later among the works it cites.
The inverse variance–flatness relation in stochastic gradient descent is critical for finding flat minima
Feng, Y. and Tu, Y · 2021
Later among the works it cites.
Wide neural networks forget less catastrophically
Mirzadeh, S. I., Chaudhry, A., Hu, H., Pascanu, R., Gorur, D., and Farajtabar, M · 2021
Later among the works it cites.
Gradient starvation: A learning proclivity in neural networks
Pezeshki, M., Kaba, O., Bengio, Y., Courville, A. C., Precup, D., and Lajoie, G · 2021
Later among the works it cites.
The intrinsic dimension of images and its impact on learning
Pope, P., Zhu, C., Abdelkader, A., Goldblum, M., and Goldstein, T · 2021
Later among the works it cites.
Chaotic dynamics are intrinsic to neural network training with SGD
Herrmann, L., Granz, M., and Landgraf, T · 2022
Later among the works it cites.
Towards understanding grokking: An effective theory of representation learning
Liu, Z., Kitouni, O., Nolte, N. S., Michaud, E. J., Tegmark, M., and Williams, M · 2022
Later among the works it cites.
A mechanistic interpretability analysis of grokking
Nanda, N. and Lieberum, T · 2022
Later among the works it cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Power, A., Burda, Y., Edwards, H., Babuschkin, I., and Misra, V · 2022
Later among the works it cites.
The slingshot mechanism: An empirical study of adaptive optimizers and the { Grokking Phenomenon}
Thilak, V., Littwin, E., Zhai, S., Saremi, O., Paiss, R., and Susskind, J. M · 2022
Later among the works it cites.
Grokking phase transitions in learning local rules with gradient descent
Žunkovič, B. and Ilievski, E · 2022
Later among the works it cites.