Fetching the paper…
Reading the bibliography…
Learning rate schedule can significantly affect generalization performance in modern neural networks, but the reasons for this are not yet understood.
Long short-term memory
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
The break-even point on optimization trajectories of deep neural networks
Jastrzebski, S., Szymczak, M., Fort, S., Arpit, D., Tabor, J., Cho, K., and Geras, K. (2020) · 2002
Earlier work this paper cites.
The large learning rate phase of deep learning: the catapult mechanism
Lewkowycz, A., Bahri, Y., Dyer, E., Sohl-Dickstein, J., and Gur-Ari, G. (2020) · 2003
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Welling, M. and Teh, Y. W. (2011) · 2011
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Krizhevsky, A. (2014) · 2014
Earlier work this paper cites.
Deep learning
Goodfellow, I., Bengio, Y., and Courville, A. (2016) · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P. (2016) · 2016
Cited alongside, same era.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y. (2017) · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K. (2017) · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Hoffer, E., Hubara, I., and Soudry, D. (2017) · 2017
Cited alongside, same era.
Stochastic gradient descent as approximate bayesian inference
Mandt, S., Hoffman, M. D., and Blei, D. M. (2017) · 2017
Cited alongside, same era.
A bayesian perspective on generalization and stochastic gradient descent
Smith, S. L. and Le, Q. V. (2017) · 2017
Later among the works it cites.
On the relation between the sharpest directions of dnn loss and the sgd step length
Jastrzebski, S., Kenton, Z., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A. (2018) · 2018
Later among the works it cites.
An empirical model of large-batch training
McCandlish, S., Kaplan, J., Amodei, D., and Team, O. D. (2018) · 2018
Later among the works it cites.
Xing, C., Arpit, D., Tsirigotis, C., and Bengio, Y. (2018) · 2018
Later among the works it cites.
Towards explaining the regularization effect of initial large learning rate in training neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Don’t decay the learning rate, increase the batch size
Smith, S. L., Kindermans, P.-J., Ying, C., and Le, Q. V. (2017) · 2017
Cited alongside, same era.
Li, Y., Wei, C., and Ma, T. (2019) · 2019
Later among the works it cites.