Fetching the paper…
Reading the bibliography…
It has long been argued that minibatch stochastic gradient descent can generalize better than large batch gradient descent in deep neural networks.
Some methods of speeding up the convergence of iteration methods
Polyak, B. T · 1964
Earlier work this paper cites.
Handbook of stochastic methods , volume 3
Gardiner, C. W. et al · 1985
Earlier work this paper cites.
On-line learning processes in artificial neural networks
Heskes, T. M. and Kappen, B · 1993
Earlier work this paper cites.
Momentum and optimal stochastic search
Orr, G. B. and Leen, T. K · 1994
Earlier work this paper cites.
Early stopping-but when?
Prechelt, L · 1998
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Qian, N · 1999
Earlier work this paper cites.
Overfitting in neural nets: Backpropagation, conjugate gradient, and early stopping
Caruana, R., Lawrence, S., and Giles, C. L · 2001
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Bottou, L · 2010
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Welling, M. and Teh, Y. W · 2011
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence rate for finite training sets
Le Roux, N., Schmidt, M., and Bach, F · 2012
Earlier work this paper cites.
Efficient backprop
LeCun, Y. A., Bottou, L., Orr, G. B., and Müller, K.-R · 2012
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87
Nesterov, Y · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G · 2013
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Krizhevsky, A · 2014
Earlier work this paper cites.
Recurrent neural network regularization
Zaremba, W., Sutskever, I., and Vinyals, O · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Cited alongside, same era.
On the influence of momentum acceleration on online learning
Yuan, K., Ying, B., and Sayed, A. H · 2016
Cited alongside, same era.
Zagoruyko, S. and Komodakis, N · 2016
Cited alongside, same era.
Understanding batch normalization
Bjorck, N., Gomes, C. P., Selman, B., and Weinberger, K. Q · 2018
Later among the works it cites.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Chaudhari, P. and Soatto, S · 2018
Later among the works it cites.
On the insufficiency of existing momentum schemes for stochastic optimization
Kidambi, R., Netrapalli, P., Jain, P., and Kakade, S · 2018
Later among the works it cites.
An empirical model of large-batch training
McCandlish, S., Kaplan, J., Amodei, D., and Team, O. D · 2018
Later among the works it cites.
How does batch normalization help optimization?
Santurkar, S., Tsipras, D., Ilyas, A., and Madry, A · 2018
Later among the works it cites.
Measuring the effects of data parallelism on neural network training
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Automated inference with adaptive batches
De, S., Yadav, A., Jacobs, D., and Goldstein, T · 2017
Cited alongside, same era.
Why momentum really works
Goh, G · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Hoffer, E., Hubara, I., and Soudry, D · 2017
Cited alongside, same era.
Three factors influencing minima in sgd
Jastrzębski, S., Kenton, Z., Arpit, D., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A · 2017
Cited alongside, same era.
Stochastic modified equations and adaptive stochastic gradient algorithms
Li, Q., Tai, C., et al · 2017
Cited alongside, same era.
Ma, S., Bassily, R., and Belkin, M · 2017
Cited alongside, same era.
Shallue, C. J., Lee, J., Antognini, J., Sohl-Dickstein, J., Frostig, R., and Dahl, G. E · 2018
Later among the works it cites.
The step decay schedule: A near optimal, geometrically decaying learning rate procedure
Ge, R., Kakade, S. M., Kidambi, R., and Netrapalli, P · 2019
Later among the works it cites.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Li, Y., Wei, C., and Ma, T · 2019
Later among the works it cites.
The effect of network width on stochastic gradient descent and generalization: an empirical study
Park, D. S., Sohl-Dickstein, J., Le, Q. V., and Smith, S. L · 2019
Later among the works it cites.
Sankararaman, K. A., De, S., Xu, Z., Huang, W. R., and Goldstein, T · 2019
Later among the works it cites.
A tail-index analysis of stochastic gradient noise in deep neural networks
Simsekli, U., Sagun, L., and Gurbuzbalaban, M · 2019
Later among the works it cites.
Momentum enables large batch training
Smith, S. L., Elsen, E., and De, S · 2019
Later among the works it cites.
Which algorithmic choices matter at which batch sizes? insights from a noisy quadratic model
Zhang, G., Li, L., Nado, Z., Martens, J., Sachdeva, S., Dahl, G. E., Shallue, C. J., and Grosse, R · 2019
Later among the works it cites.
Batch normalization biases residual blocks towards the identity function in deep networks
De, S. and Smith, S. L · 2020
Closest in time.