Fetching the paper…
Reading the bibliography…
This study investigates how weight decay affects the update behavior of individual neurons in deep neural networks through a combination of applied analysis and experimentation.
Qiao, S., Wang, H., Liu, C., Shen, W., and Yuille, A. L · 1903
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 1904
Earlier work this paper cites.
Large batch optimization for deep learning: Training bert in 76 minutes
You, Y., Li, J., Reddi, S., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., Keutzer, K., and Hsieh, C.-J · 1904
Earlier work this paper cites.
Online normalization for training neural networks
Chiley, V., Sharapov, I., Kosson, A., Koster, U., Reece, R., Samaniego de la Fuente, S., Subbiah, V., and James, M · 1905
Earlier work this paper cites.
An exponential learning rate schedule for deep learning
Li, Z. and Arora, S · 1910
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 1912
Earlier work this paper cites.
Adamp: Slowing down the slowdown for momentum optimizers on scale-invariant weights
Heo, B., Chun, S., Oh, S. J., Han, D., Yun, S., Kim, G., Uh, Y., and Ha, J.-W · 2006
Earlier work this paper cites.
A spherical analysis of adam with batch normalization
Roburin, S., de Mont-Marin, Y., Bursuc, A., Marlet, R., Perez, P., and Aubry, M · 2006
Earlier work this paper cites.
Spherical motion dynamics: Learning dynamics of normalized neural network using sgd and weight decay
Wan, R., Zhu, Z., Zhang, X., and Sun, J · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Reconciling modern deep learning with traditional optimization analyses: The intrinsic learning rate
Li, Z., Lyu, K., and Arora, S · 2010
Earlier work this paper cites.
On the overlooked pitfalls of weight decay and how to mitigate them: A gradient-norm perspective
Xie, Z., zhiqiang xu, Zhang, J., Sato, I., and Sugiyama, M · 2011
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jegou, H · 2012
Earlier work this paper cites.
Report on the 11th IWSLT evaluation campaign
Cettolo, M., Niehues, J., Stüker, S., Bentivogli, L., and Federico, M · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2015
Earlier work this paper cites.
Path-sgd: Path-normalized optimization in deep neural networks
Neyshabur, B., Salakhutdinov, R. R., and Srebro, N · 2015
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2015
Cited alongside, same era.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Cited alongside, same era.
Deep Learning
Goodfellow, I., Bengio, Y., and Courville, A · 2016
Measuring the effects of data parallelism on neural network training
Shallue, C. J., Lee, J., Antognini, J., Sohl-Dickstein, J., Frostig, R., and Dahl, G. E · 2019
Later among the works it cites.
Pytorch image models
Wightman, R · 2019
Later among the works it cites.
Three mechanisms of weight decay regularization
Zhang, G., Wang, C., Xu, B., and Grosse, R · 2019
Later among the works it cites.
On the validity of modeling SGD with stochastic differential equations (SDEs)
Li, Z., Malladi, S., and Arora, S · 2021
Later among the works it cites.
Learning by turning: Neural architecture aware optimisation
Liu, Y., Bernstein, J., Meister, M., and Yue, Y · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, T. and Kingma, D. P · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2017
Cited alongside, same era.
L2 regularization versus batch and weight normalization
Van Laarhoven, T · 2017
Cited alongside, same era.
Large batch training of convolutional networks
You, Y., Gitman, I., and Ginsburg, B · 2017
Cited alongside, same era.
Norm matters: efficient and accurate normalization schemes in deep networks
Hoffer, E., Banner, R., Golan, I., and Soudry, D · 2018
Cited alongside, same era.
Jia, X., Song, S., He, W., Wang, Y., Rong, H., Zhou, F., Xie, L., Guo, Z., Yang, Y., Yu, L., et al · 2018
Cited alongside, same era.
Later among the works it cites.
Fixnorm: Dissecting weight decay for training deep neural networks
Zhou, Y., Sun, Y., and Zhong, Z · 2021
Later among the works it cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Later among the works it cites.
Training scale-invariant neural networks on the sphere can happen in three regimes
Kodryan, M., Lobacheva, E., Nakhodnov, M., and Vetrov, D. P · 2022
Later among the works it cites.
Fast mixing of stochastic gradient descent with normalization and weight decay
Li, Z., Wang, T., and Yu, D · 2022
Later among the works it cites.
On the SDEs and scaling rules for adaptive gradient algorithms
Malladi, S., Lyu, K., Panigrahi, A., and Arora, S · 2022
Later among the works it cites.
Why do we need weight decay in modern deep learning?
Andriushchenko, M., D’Angelo, F., Varre, A., and Flammarion, N · 2023
Closest in time.
Symbolic discovery of optimization algorithms
Chen, X., Liang, C., Huang, D., Real, E., Wang, K., Pham, H., Dong, X., Luong, T., Hsieh, C.-J., Lu, Y., and Le, Q. V · 2023
Closest in time.
When and why momentum accelerates sgd: An empirical study
Fu, J., Wang, B., Zhang, H., Zhang, Z., Chen, W., and Zheng, N · 2023
Closest in time.
Analyzing and improving the training dynamics of diffusion models
Karras, T., Aittala, M., Lehtinen, J., Hellsten, J., Aila, T., and Laine, S · 2023
Closest in time.
llm-baseline
Pagliardini, M · 2023
Closest in time.
Ghost noise for regularizing deep neural networks
Kosson, A., Fan, D., and Jaggi, M · 2024
Closest in time.