Fetching the paper…
Reading the bibliography…
Large-batch training approaches have enabled researchers to utilize large-scale distributed processing and greatly accelerate deep-neural net (DNN) training.
A stochastic approximation method
H. Robbins and S. Monro · 1985
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
M. P. Marcus, M. A. Marcinkiewicz, and B. Santorini · 1993
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
N. Qian · 1999
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
M. D. Zeiler · 2012
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
A. Krizhevsky · 2014
Earlier work this paper cites.
Efficient mini-batch training for stochastic optimization
M. Li, T. Zhang, Y. Chen, and A. J. Smola · 2014
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
J. Martens and R. Grosse · 2015
Earlier work this paper cites.
Multi-scale context aggregation by dilated convolutions
F. Yu and V. Koltun · 2015
Cited alongside, same era.
Deep learning
I. Goodfellow, Y. Bengio, A. Courville, and Y. Bengio · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Firecaffe: near-linear acceleration of deep neural network training on compute clusters
F. N. Iandola, M. W. Moskewicz, K. Ashraf, and K. Keutzer · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2016
Cited alongside, same era.
E. Hoffer, I. Hubara, and D. Soudry · 2017
Later among the works it cites.
Scaling Distributed Machine Learning with System and Algorithm Co-design
M. Li · 2017
Later among the works it cites.
Neural machine translation (seq2seq) tutorial
M. Luong, E. Brevdo, and R. Zhao · 2017
Later among the works it cites.
P. Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaev, G. Venkatesh, et al · 2017
Later among the works it cites.
Don’t decay the learning rate, increase the batch size
S. L. Smith, P.-J. Kindermans, and Q. V. Le · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Neil, M. Pfeiffer, and S.-C. Liu · 2016
Cited alongside, same era.
Recurrent residual learning for sequence classification
Y. Wang and F. Tian · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, et al · 2016
Cited alongside, same era.
Extremely large minibatch sgd: Training resnet-50 on imagenet in 15 minutes
T. Akiba, S. Suzuki, and K. Fukuda · 2017
Cited alongside, same era.
V. Codreanu, D. Podareanu, and V. Saletore · 2017
Cited alongside, same era.
Adabatch: Adaptive batch sizes for training deep neural networks
A. Devarakonda, M. Naumov, and M. Garland · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Cited alongside, same era.
Scaling sgd batch size to 32k for imagenet training
Y. You, I. Gitman, and B. Ginsburg · 2017
Later among the works it cites.
Y. You, Z. Zhang, C. Hsieh, J. Demmel, and K. Keutzer · 2017
Later among the works it cites.
X. Jia, S. Song, W. He, Y. Wang, H. Rong, F. Zhou, L. Xie, Z. Guo, Y. Yang, L. Yu, et al · 2018
Later among the works it cites.
Second-order optimization method for large mini-batch: Training resnet-50 on imagenet in 35 epochs
K. Osawa, Y. Tsuji, Y. Ueno, A. Naruse, R. Yokota, and S. Matsuoka · 2018
Later among the works it cites.
Measuring the effects of data parallelism on neural network training
C. J. Shallue, J. Lee, J. Antognini, J. Sohl-Dickstein, R. Frostig, and G. E. Dahl · 2018
Later among the works it cites.
Image classification at supercomputer scale
C. Ying, S. Kumar, D. Chen, T. Wang, and Y. Cheng · 2018
Later among the works it cites.