Fetching the paper…
Reading the bibliography…
Weight decay is a broadly used technique for training state-of-the-art deep networks from image classification to large language models.
On the stability of inverse problems
Tikhonov, A. N · 1943
Earlier work this paper cites.
An application of the wiener-kolmogorov smoothing theory to matrix inversion
Foster, M · 1961
Earlier work this paper cites.
Application of ridge analysis to regression problems
Hoerl, A. R · 1962
Earlier work this paper cites.
Ridge regression: Biased estimation for nonorthogonal problems
Hoerl, A. E. and Kennard, R. W · 1970
Earlier work this paper cites.
A simple weight decay can improve generalization
Krogh, A. and Hertz, J · 1991
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Moulines, E. and Bach, F · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S. and Ben-David, S · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Earlier work this paper cites.
Snapshot ensembles: Train 1, get m for free
Huang, G., Li, Y., Pleiss, G., Liu, Z., Hopcroft, J. E., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Three factors influencing minima in sgd
Jastrzebski, S., Kenton, Z., Arpit, D., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A · 2017
Earlier work this paper cites.
Empirical analysis of the hessian of over-parametrized neural networks
Sagun, L., Evci, U., Guney, V. U., Dauphin, Y., and Bottou, L · 2017
Earlier work this paper cites.
L2 regularization versus batch and weight normalization
Van Laarhoven, T · 2017
Earlier work this paper cites.
Tiny imagenet challenge
Wu, J., Zhang, Q., and Xu, G · 2017
Earlier work this paper cites.
Dissecting adam: The sign, magnitude and variance of stochastic gradients
Balles, L. and Hennig, P · 2018
Earlier work this paper cites.
Norm matters: efficient and accurate normalization schemes in deep networks
Hoffer, E., Banner, R., Golan, I., and Soudry, D · 2018
Cited alongside, same era.
The full spectrum of deepnet hessians at scale: Dynamics with sgd training and sample size
Papyan, V · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N · 2018
Cited alongside, same era.
Three mechanisms of weight decay regularization
Zhang, G., Wang, C., Xu, B., and Grosse, R · 2018
Cited alongside, same era.
Zhu, Z., Wu, J., Yu, B., Wu, L., and Ma, J · 2018
Cited alongside, same era.
Xie, Z., Sato, I., and Sugiyama, M · 2020
Later among the works it cites.
Pyhessian: Neural networks through the lens of the hessian
Yao, Z., Gholami, A., Keutzer, K., and Mahoney, M. W · 2020
Later among the works it cites.
Understanding decoupled and early weight decay
Bjorck, J., Weinberger, K. Q., and Gomes, C · 2021
Later among the works it cites.
Gradient descent on neural networks typically occurs at the edge of stability
Cohen, J. M., Kaur, S., Li, Y., Kolter, J. Z., and Talwalkar, A · 2021
Later among the works it cites.
Label noise sgd provably prefers flat global minimizers
Damian, A., Ma, T., and Lee, J. D · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Openwebtext corpus, 2019
Gokaslan, A., Cohen, V., Pavlick, E., and Tellex, S · 2019
Cited alongside, same era.
The implicit bias of gradient descent on nonseparable data
Ji, Z. and Telgarsky, M · 2019
Cited alongside, same era.
A study of bfloat16 for deep learning training
Kalamkar, D., Mudigere, D., Mellempudi, N., Das, D., Banerjee, K., Avancha, S., Vooturi, D. T., Jammalamadaka, N., Huang, J., Yuen, H., et al · 2019
Cited alongside, same era.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Li, Y., Wei, C., and Ma, T · 2019
Cited alongside, same era.
An exponential learning rate schedule for deep learning
Li, Z. and Arora, S · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Cited alongside, same era.
Lewkowycz, A · 2021
Later among the works it cites.
Stochastic gradient descent with noise of machine learning type. Part I: Discrete time analysis
Wojtowytsch, S · 2021
Later among the works it cites.
Adaptive gradient methods at the edge of stability
Cohen, J. M., Ghorbani, B., Krishnan, S., Agarwal, N., Medapati, S., Badura, M., Suo, D., Cardoze, D., Nado, Z., Dahl, G. E., et al · 2022
Later among the works it cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Later among the works it cites.
Training scale-invariant neural networks on the sphere can happen in three regimes
Kodryan, M., Lobacheva, E., Nakhodnov, M., and Vetrov, D. P · 2022
Later among the works it cites.
Power-law escape rate of SGD
Mori, T., Liu, Z., Liu, K., and Ueda, M · 2022
Later among the works it cites.
Label noise (stochastic) gradient descent implicitly solves the lasso for quadratic parametrisation
Pillaud-Vivien, L., Reygner, J., and Flammarion, N · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
Scao, T. L., Fan, A., Akiki, C., Pavlick, E., Ilić, S., Hesslow, D., Castagné, R., Luccioni, A. S., Yvon, F., Gallé, M., et al · 2022
Later among the works it cites.
SGD with large step sizes learns sparse features
Andriushchenko, M., Varre, A., Pillaud-Vivien, L., and Flammarion, N · 2023
Closest in time.
Nanogpt repository, 2023
Karpathy, A · 2023
Closest in time.
Rotational equilibrium: How weight decay balances learning across neural networks
Kosson, A., Messmer, B., and Jaggi, M · 2023
Closest in time.
Understanding the effectiveness of early weight averaging for training large language models
Sanyal, S., Kaddour, J., Kumar, A., and Sanghavi, S · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Closest in time.