Fetching the paper…
Reading the bibliography…
Large-scale distributed training of deep neural networks suffer from the generalization gap caused by the increase in the effective mini-batch size.
Natural gradient works efficiently in learning
S.-I. Amari · 1998
Earlier work this paper cites.
Topmoumoute online natural gradient algorithm
N. Le Roux, P.-A. Manzagol, and Y. Bengio · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Deep learning via Hessian-free optimization
J. Martens · 2010
Earlier work this paper cites.
Batch Normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Optimizing neural networks with Kronecker-factored approximate curvature
J. Martens and R. Grosse · 2015
Earlier work this paper cites.
Chainer: a next-generation open source framework for deep learning
S. Tokui, K. Oono, S. Hido, and J. Clayton · 2015
Earlier work this paper cites.
TensorFlow: A system for large-scale machine learning
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, J. Levenberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P. Tucker, V. Vasudevan, P. Warden, M. Wicke, Y. Yu, and X. Zheng · 2016
Earlier work this paper cites.
A Kronecker-factored approximate Fisher matrix for convolution layers
R. Grosse and J. Martens · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
ChainerMN: Scalable distributed deep learning framework
T. Akiba, K. Fukuda, and S. Suzuki · 2017
Cited alongside, same era.
Extremely large minibatch sgd: Training ResNet-50 on ImageNet in 15 minutes
T. Akiba, S. Suzuki, and K. Fukuda · 2017
Cited alongside, same era.
Distributed second-order optimization using Kronecker-factored approximations
J. Ba, R. Grosse, and J. Martens · 2017
Cited alongside, same era.
Practical Gauss-Newton optimisation for deep learning
A. Botev, H. Ritter, and D. Barber · 2017
Cited alongside, same era.
AdaBatch: Adaptive batch sizes for training deep neural networks
A. Devarakonda, M. Naumov, and M. Garland · 2017
Cited alongside, same era.
X. Jia, S. Song, W. He, Y. Wang, H. Rong, F. Zhou, L. Xie, Z. Guo, Y. Yang, L. Yu, T. Chen, G. Hu, S. Shi, and X. Chu · 2018
Closest in time.
Universal statistics of Fisher information in deep neural networks: Mean field approach
R. Karakida, S. Akaho, and S.-i. Amari · 2018
Closest in time.
Don’t use large mini-batches, use local SGD
T. Lin, S. U. Stich, and M. Jaggi · 2018
Closest in time.
Konecker-factored curvature approximations for recurrent neural networks
J. Martens, J. Ba, and M. Johnson · 2018
Closest in time.
Massively distributed SGD: ImageNet/ResNet-50 training in a flash
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Goyal, P. Dollar, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Cited alongside, same era.
Train longer, generalize better: Closing the generalization gap in large batch training of neural networks
E. Hoffer, I. Hubara, and D. Soudry · 2017
Cited alongside, same era.
L2 regularization versus batch and weight normalization
T. van Laarhoven · 2017
Cited alongside, same era.
Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation
Y. Wu, E. Mansimov, S. Liao, R. Grosse, and J. Ba · 2017
Cited alongside, same era.
MixUp as locally linear out-of-manifold regularization
H. Guo, Y. Mao, and R. Zhang · 2018
Cited alongside, same era.
Image classification at supercomputer scale
C. Ying, S. Kumar, D. Chen, T. Wang, and Y. Cheng
Cited in the paper.
H. Mikami, H. Suganuma, P. U-chupala, Y. Tanaka, and Y. Kageyama · 2018
Closest in time.
Measuring the effects of data parallelism on neural network training
C. J. Shallue, J. Lee, J. Antognini, J. Sohl-Dickstein, R. Frostig, and G. E. Dahl · 2018
Closest in time.
Don’t decay the learning rate, increase the batch size
S. L. Smith, P.-J. Kindermans, and Q. V. Le · 2018
Closest in time.
ImageNet training in minutes
Y. You, Z. Zhang, C.-J. Hsieh, J. Demmel, and K. Keutzer · 2018
Closest in time.
Noisy natural gradient as variational inference
G. Zhang, S. Sun, D. Duvenaud, and R. Grosse · 2018
Closest in time.
Mixup: Beyond empirical risk minimization
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz · 2018
Closest in time.