Fetching the paper…
Reading the bibliography…
In stochastic optimization, using large batch sizes during training can leverage parallel resources to produce faster wall-clock training times per training epoch.
S.-I. Amari, “Natural gradient works efficiently in learning,”
1998
Earlier work this paper cites.
A. Krizhevsky and G. Hinton, “Learning multiple layers of features from tiny images,” tech. rep., Citeseer, 2009
2009
Earlier work this paper cites.
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng, “Reading digits in natural images with unsupervised feature learning,” 2011
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in
2012
Earlier work this paper cites.
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Earlier work this paper cites.
R. Grosse and J. Martens, “A Kronecker-factored Approximate Fisher matrix for convolution layers,” in
2016
Earlier work this paper cites.
F. Roosta-Khorasani and M. W. Mahoney, “Sub-sampled Newton methods I: Globally convergent algorithms,” tech. rep., 2016 · 2016
Earlier work this paper cites.
F. Roosta-Khorasani and M. W. Mahoney, “Sub-sampled Newton methods II: Local convergence rates,” tech. rep., 2016 · 2016
Earlier work this paper cites.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Y. You, Z. Zhang, C. Hsieh, and J. Demmel, “100-epoch imagenet training with alexnet in 24 minutes,”
2017
Cited alongside, same era.
J. Ba, R. Grosse, and J. Martens, “Distributed second-order optimization using Kronecker-factored Approximations,” in
S. McCandlish, J. Kaplan, D. Amodei, and O. D. Team, “An empirical model of large-batch training,”
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
2017
Cited alongside, same era.
E. Hoffer, I. Hubara, and D. Soudry, “Train longer, generalize better: Closing the generalization gap in large batch training of neural networks,” in
2017
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Later among the works it cites.
B. Ginsburg, I. Gitman, and Y. You, “Large batch training of convolutional networks with layer-wise adaptive rate scaling,” 2018
2018
Later among the works it cites.
X. Jia, S. Song, W. He, Y. Wang, H. Rong, F. Zhou, L. Xie, Z. Guo, Y. Yang, L. Yu,
2018
Later among the works it cites.
C. H. Martin and M. W. Mahoney, “Traditional and heavy-tailed self regularization in neural network models,” in
2019
Closest in time.
G. Zhang, C. Wang, B. Xu, and R. Grosse, “Three mechanisms of weight decay regularization,” in
2019
Closest in time.