Fetching the paper…
Reading the bibliography…
The success of SGD in deep learning has been ascribed by prior works to the implicit bias induced by finite batch sizes ("SGD noise").
SGD on neural networks learns functions of increasing complexity
Nakkiran, P., Kaplun, G., Kalimeris, D., Yang, T., Edelman, B. L., Zhang, F., and Barak, B · 1905
Earlier work this paper cites.
Deep double descent: Where bigger models and more data hurt
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I · 1912
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Efficient BackProp , pp. 9–50
LeCun, Y., Bottou, L., Orr, G. B., and Müller, K. R · 1998
Earlier work this paper cites.
The break-even point on optimization trajectories of deep neural networks
Jastrzebski, S., Szymczak, M., Fort, S., Arpit, D., Tabor, J., Cho, K., and Geras, K. J · 2002
Earlier work this paper cites.
The large learning rate phase of deep learning: the catapult mechanism
Lewkowycz, A., Bahri, Y., Dyer, E., Sohl-Dickstein, J., and Gur-Ari, G · 2003
Earlier work this paper cites.
The pitfalls of simplicity bias in neural networks
Shah, H., Tamuly, K., Raghunathan, A., Jain, P., and Netrapalli, P · 2006
Earlier work this paper cites.
Catastrophic fisher explosion: Early phase fisher matrix impacts generalization
Jastrzebski, S., Arpit, D., Åstrand, O., Kerg, G., Wang, H., Xiong, C., Socher, R., Cho, K., and Geras, K. J · 2012
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Earlier work this paper cites.
Accurate, large minibatch SGD: training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R. B., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Earlier work this paper cites.
Three factors influencing minima in SGD
Jastrzebski, S., Kenton, Z., Arpit, D., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A. J · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks, 2018
Hoffer, E., Hubara, I., and Soudry, D · 2018
Earlier work this paper cites.
An alternative view: When does SGD escape local minima?
Kleinberg, B., Li, Y., and Yuan, Y · 2018
Earlier work this paper cites.
Revisiting small batch training for deep neural networks
Masters, D. and Luschi, C · 2018
Earlier work this paper cites.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Wu, L., Ma, C., and E, W · 2018
Earlier work this paper cites.
A walk with sgd, 2018
Xing, C., Arpit, D., Tsirigotis, C., and Bengio, Y · 2018
Earlier work this paper cites.
Control batch size and learning rate to generalize well: Theoretical and empirical evidence
He, F., Liu, T., and Tao, D · 2019
Cited alongside, same era.
A recipe for training neural networks, 2019
Karpathy, A · 2019
Cited alongside, same era.
Komatsuzaki, A · 2019
Cited alongside, same era.
Similarity of neural network representations revisited
Kornblith, S., Norouzi, M., Lee, H., and Hinton, G. E · 2019
Cited alongside, same era.
The implicit regularization of stochastic gradient flow for least squares, 2020
Ali, A., Dobriban, E., and Tibshirani, R. J · 2020
Cited alongside, same era.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
The three stages of learning dynamics in high-dimensional kernel methods, 2021
Ghosh, N., Mei, S., and Yu, B · 2021
Later among the works it cites.
Shape matters: Understanding the implicit bias of the noise covariance
HaoChen, J. Z., Wei, C., Lee, J., and Ma, T · 2021
Later among the works it cites.
A large batch optimizer reality check: Traditional, generic optimizers suffice across batch sizes
Nado, Z., Gilmer, J., Shallue, C. J., Anil, R., and Dahl, G. E · 2021
Later among the works it cites.
The deep bootstrap framework: Good online learners are good offline generalizers
Nakkiran, P., Neyshabur, B., and Sedghi, H · 2021
Later among the works it cites.
On the origin of implicit regularization in stochastic gradient descent
Smith, S. L., Dherin, B., Barrett, D. G. T., and De, S · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Blanc, G., Gupta, N., Valiant, G., and Valiant, P · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the neural tangent kernel
Fort, S., Dziugaite, G. K., Paul, M., Kharaghani, S., Roy, D. M., and Ganguli, S · 2020
Cited alongside, same era.
Linear mode connectivity and the lottery ticket hypothesis
Frankle, J., Dziugaite, G. K., Roy, D., and Carbin, M · 2020
Cited alongside, same era.
Let’s agree to agree: Neural networks share classification order on real datasets
Hacohen, G., Choshen, L., and Weinshall, D · 2020
Cited alongside, same era.
On the generalization benefit of noise in stochastic gradient descent
Smith, S. L., Elsen, E., and De, S · 2020
Cited alongside, same era.
Deep learning through the lens of example difficulty
Baldock, R., Maennel, H., and Neyshabur, B · 2021
Cited alongside, same era.
Sgd with large step sizes learns sparse features, 2022
Andriushchenko, M., Varre, A., Pillaud-Vivien, L., and Flammarion, N · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Later among the works it cites.
Scaling laws and interpretability of learning from repeated data
Hernandez, D., Brown, T., Conerly, T., DasSarma, N., Drain, D., El-Showk, S., Elhage, N., Hatfield-Dodds, Z., Henighan, T., Hume, T., et al · 2022
Later among the works it cites.
Achieving small-batch accuracy with large-batch scalability via adaptive learning rate adjustment, 2022
Lee, S. and Avestimehr, S · 2022
Later among the works it cites.
Liu, Z., Mao, H., Wu, C., Feichtenhofer, C., Darrell, T., and Xie, S · 2022
Later among the works it cites.
Implicit bias of the step size in linear diagonal neural networks
Nacson, M. S., Ravichandran, K., Srebro, N., and Soudry, D · 2022
Later among the works it cites.
Disentangling the mechanisms behind implicit regularization in sgd, 2022
Novack, Z., Kaur, S., Marwah, T., Garg, S., and Lipton, Z. C · 2022
Later among the works it cites.
In-context learning and induction heads, 2022
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2022
Later among the works it cites.
Implicit regularization or implicit conditioning? exact risk trajectories of SGD in high dimensions
Paquette, C., Paquette, E., Adlam, B., and Pennington, J · 2022
Later among the works it cites.
To repeat or not to repeat: Insights from scaling llm under token-crisis, 2023
Xue, F., Fu, Y., Zhou, W., Zheng, Z., and You, Y · 2023
Closest in time.