Fetching the paper…
Reading the bibliography…
Large-scale distributed training of deep neural networks results in models with worse generalization performance as a result of the increase in the effective mini-batch size.
S.-i. Amari, “Natural Gradient Works Efficiently in Learning,” Neural Computation , vol. 10, pp. 251–276, 1998
1998
Earlier work this paper cites.
N. L. Roux, P.-a. Manzagol, and Y. Bengio, “Topmoumoute Online Natural Gradient Algorithm,” in Advances in Neural Information Processing Systems 20 , 2008, pp. 849–856
2008
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A Large-Scale Hierarchical Image Database,” in IEEE Conference on Computer Vision and Pattern Recognition , 2009, pp. 248–255
2009
Earlier work this paper cites.
J. Martens, “Deep learning via Hessian-free optimization,” in Proceedings of the 27th International Conference on Machine Learning , 2010, pp. 735–742
2010
Earlier work this paper cites.
S. Tokui, H. Yamazaki Vincent, R. Okuta, T. Akiba, Y. Niitani, T. Ogawa, S. Saito, S. Suzuki, K. Uenishi, and B. Vogel, “Chainer: A Deep Learning Framework for Accelerating the Research Cycle,” in Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , 2019, pp. 2002–2011
2011
Earlier work this paper cites.
J. Martens and R. Grosse, “Optimizing Neural Networks with Kronecker-factored Approximate Curvature,” in Proceedings of the 32nd International Conference on International Conference on Machine Learning , 2015, pp. 2408–2417
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift,” Proceedings of the 32nd International Conference on International Conference on Machine Learning , pp. 448–456, 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
R. Grosse and J. Martens, “A Kronecker-factored approximate Fisher matrix for convolution layers,” in Proceedings of the 33rd International Conference on International Conference on Machine Learning , 2016, pp. 573–582
2016
Earlier work this paper cites.
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, J. Levenberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P. Tucker, V. Vasudevan, P. Warden, M. Wicke, Y. Yu, and X. Zheng, “TensorFlow: A system for large-scale machine learning,” in 12th USENIX Symposium on Operating Systems Design and Implementation , 2016, pp. 265–283
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. Ba, R. Grosse, and J. Martens, “Distributed second-order optimization using Kronecker-factored approximations,” in International Conference on Learning Representations , 2017
2017
Cited alongside, same era.
Y. Wu, E. Mansimov, S. Liao, R. Grosse, and J. Ba, “Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation,” in Advances in Neural Information Processing Systems , 2017, pp. 5279–5288
2017
Cited alongside, same era.
A. Botev, H. Ritter, and D. Barber, “Practical Gauss-Newton Optimisation for Deep Learning,” in Proceedings of the 34th International Conference on Machine Learning , 2017, pp. 557–565
2017
Cited alongside, same era.
2017
Cited alongside, same era.
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “Mixup: Beyond Empirical Risk Minimization,” in International Conference on Learning Representations , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
C. J. Shallue, J. Lee, J. Antognini, J. Sohl-Dickstein, R. Frostig, and G. E. Dahl, “Measuring the Effects of Data Parallelism on Neural Network Training,” Journal of Machine Learning Research , vol. 20, no. 112, pp. 1–49, 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Hoffer, R. Banner, I. Golan, and D. Soudry, “Norm matters: Efficient and accurate normalization schemes in deep networks,” in Advances in Neural Information Processing Systems , 2018, pp. 2160–2170
2018
Cited alongside, same era.
S. L. Smith, P.-J. Kindermans, C. Ying, and Q. V. Le, “Don’t Decay the Learning Rate, Increase the Batch Size,” in International Conference on Learning Representations , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
J. Martens and M. Johnson, “Kronecker-factored Curvature Approximations for Recurrent Neural Networks,” in International Conference on Learning Representations , 2018
2018
Cited alongside, same era.
G. Zhang, S. Sun, D. Duvenaud, and R. Grosse, “Noisy Natural Gradient as Variational Inference,” in Proceedings of the 35th International Conference on Machine Learning , 2018
2018
Cited alongside, same era.
2019
Later among the works it cites.
K. Osawa, Y. Tsuji, Y. Ueno, A. Naruse, R. Yokota, and S. Matsuoka, “Large-Scale Distributed Second-Order Optimization Using Kronecker-Factored Approximate Curvature for Deep Convolutional Neural Networks,” in IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 12 359–12 367
2019
Later among the works it cites.
Y. Tsuji, K. Osawa, Y. Ueno, A. Naruse, R. Yokota, and S. Matsuoka, “Performance Optimizations and Analysis of Distributed Deep Learning with Approximated Second-Order Optimization Method,” in Proceedings of the 48th International Conference on Parallel Processing: Workshops , 2019, pp. 21:1–21:8
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
S.-i. Amari, R. Karakida, and M. Oizumi, “Fisher Information and Natural Gradient Learning in Random Deep Networks,” in The 22nd International Conference on Artificial Intelligence and Statistics , 2019, pp. 694–702
2019
Later among the works it cites.
Y. Ueno and R. Yokota, “Exhaustive Study of Hierarchical AllReduce Patterns for Large Messages Between GPUs,” in 19th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID) , 2019, pp. 430–439
2019
Later among the works it cites.