Fetching the paper…
Reading the bibliography…
A fundamental property of deep learning normalization techniques, such as batch normalization, is making the pre-normalization parameters scale invariant.
Wide-minima density hypothesis and the explore-exploit learning rate schedule
Iyer, N., Thejas, V., Kwatra, N., Ramjee, R., and Sivathanu, M. (2020) · 2003
Earlier work this paper cites.
The large learning rate phase of deep learning: the catapult mechanism
Lewkowycz, A., Bahri, Y., Dyer, E., Sohl-Dickstein, J., and Gur-Ari, G. (2020) · 2003
Earlier work this paper cites.
A spherical analysis of adam with batch normalization
Roburin, S., de Mont-Marin, Y., Bursuc, A., Marlet, R., Pérez, P., and Aubry, M. (2020) · 2006
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research
Deng, L. (2012) · 2012
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C. (2015) · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2015) · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Zhang, X., Zhao, J., and LeCun, Y. (2015) · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E. (2016) · 2016
Earlier work this paper cites.
Riemannian approach to batch normalization
Cho, M. and Lee, J. (2017) · 2017
Earlier work this paper cites.
Three factors influencing minima in sgd
Jastrzebski, S., Kenton, Z., Arpit, D., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A. (2017) · 2017
Earlier work this paper cites.
L2 regularization versus batch and weight normalization
Van Laarhoven, T. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Norm matters: efficient and accurate normalization schemes in deep networks
Hoffer, E., Banner, R., Golan, I., and Soudry, D. (2018) · 2018
Cited alongside, same era.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Wu, L., Ma, C., et al. (2018) · 2018
Cited alongside, same era.
Xing, C., Arpit, D., Tsirigotis, C., and Bengio, Y. (2018) · 2018
Cited alongside, same era.
Theoretical analysis of auto rate-tuning by batch normalization
Arora, S., Li, Z., and Lyu, K. (2019) · 2019
Cited alongside, same era.
Online normalization for training neural networks
Chiley, V., Sharapov, I., Kosson, A., Koster, U., Reece, R., Samaniego de la Fuente, S., Subbiah, V., and James, M. (2019) · 2019
Cited alongside, same era.
Implicit gradient regularization
Barrett, D. and Dherin, B. (2021) · 2021
Later among the works it cites.
Gradient descent on neural networks typically occurs at the edge of stability
Cohen, J., Kaur, S., Li, Y., Kolter, J. Z., and Talwalkar, A. (2021) · 2021
Later among the works it cites.
Catastrophic fisher explosion: Early phase fisher matrix impacts generalization
Jastrzebski, S., Arpit, D., Astrand, O., Kerg, G. B., Wang, H., Xiong, C., Socher, R., Cho, K., and Geras, K. J. (2021) · 2021
Later among the works it cites.
HPC resources of the higher school of economics
Kostenetskiy, P. S., Chulkevich, R. A., and Kozyrev, V. I. (2021) · 2021
Later among the works it cites.
Neural mechanics: Symmetry and broken conservation laws in deep learning dynamics
Kunin, D., Sagastuy-Brena, J., Ganguli, S., Yamins, D. L., and Tanaka, H. (2021) · 2021
Later among the works it cites.
On the periodic behavior of neural network training with batch normalization and weight decay
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jastrzebski, S., Kenton, Z., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A. (2019) · 2019
Cited alongside, same era.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Li, Y., Wei, C., and Ma, T. (2019) · 2019
Cited alongside, same era.
Three mechanisms of weight decay regularization
Zhang, G., Wang, C., Xu, B., and Grosse, R. (2019) · 2019
Cited alongside, same era.
The break-even point on optimization trajectories of deep neural networks
Jastrzebski, S., Szymczak, M., Fort, S., Arpit, D., Tabor, J., Cho, K., and Geras, K. (2020) · 2020
Cited alongside, same era.
An exponential learning rate schedule for deep learning
Li, Z. and Arora, S. (2020) · 2020
Cited alongside, same era.
Deep double descent: Where bigger models and more data hurt
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I. (2020) · 2020
Cited alongside, same era.
On the interplay between noise and curvature and its effect on optimization and generalization
Thomas, V., Pedregosa, F., Merriënboer, B., Manzagol, P.-A., Bengio, Y., and Le Roux, N. (2020) · 2020
Cited alongside, same era.
Lobacheva, E., Kodryan, M., Chirkova, N., Malinin, A., and Vetrov, D. P. (2021) · 2021
Later among the works it cites.
On the origin of implicit regularization in stochastic gradient descent
Smith, S. L., Dherin, B., Barrett, D., and De, S. (2021) · 2021
Later among the works it cites.
Spherical motion dynamics: Learning dynamics of normalized neural network using sgd and weight decay
Wan, R., Zhu, Z., Zhang, X., and Sun, J. (2021) · 2021
Later among the works it cites.
A loss curvature perspective on training instabilities of deep learning models
Gilmer, J., Ghorbani, B., Garg, A., Kudugunta, S., Neyshabur, B., Cardoze, D., Dahl, G. E., Nado, Z., and Firat, O. (2022) · 2022
Closest in time.
Robust training of neural networks using scale invariant architectures
Li, Z., Bhojanapalli, S., Zaheer, M., Reddi, S., and Kumar, S. (2022) · 2022
Closest in time.
Towards practical control of singular values of convolutional layers
Senderovich, A., Bulatova, E., Obukhov, A., and Rakhuba, M. (2022) · 2022
Closest in time.