Fetching the paper…
Reading the bibliography…
Training neural networks with many processors can reduce time-to-solution; however, it is challenging to maintain convergence and efficiency at large scales.
D. C. Liu and J. Nocedal, “On the limited memory bfgs method for large scale optimization,”
1989
Earlier work this paper cites.
R. Thakur, R. Rabenseifner, and W. Gropp, “Optimization of collective communication operations in MPICH,”
2005
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in
2009
Earlier work this paper cites.
P. Patarasuk and X. Yuan, “Bandwidth optimal all-reduce algorithms for clusters of workstations,”
2009
Earlier work this paper cites.
A. Krizhevsky, G. Hinton
2009
Earlier work this paper cites.
B. Recht, C. Re, S. Wright, and F. Niu, “Hogwild: A lock-free approach to parallelizing stochastic gradient descent,” in
2011
Earlier work this paper cites.
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su, “Scaling distributed machine learning with the parameter server,” in
2014
Earlier work this paper cites.
J. Martens and R. Grosse, “Optimizing neural networks with kronecker-factored approximate curvature,” in
2015
Earlier work this paper cites.
S. Zhang, A. E. Choromanska, and Y. LeCun, “Deep learning with elastic averaging SGD,” in
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Earlier work this paper cites.
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard
2016
Earlier work this paper cites.
P. H. Jin, Q. Yuan, F. Iandola, and K. Keutzer, “How to scale distributed deep learning?”
2016
Earlier work this paper cites.
I. Mitliagkas, C. Zhang, S. Hadjis, and C. Ré, “Asynchrony begets momentum, with an application to deep learning,” in
2016
Cited alongside, same era.
R. Grosse and J. Martens, “A Kronecker-factored approximate Fisher matrix for convolution layers,”
2016
Cited alongside, same era.
2017
Cited alongside, same era.
J. Carrasquilla and R. G. Melko, “Machine learning phases of matter,”
2017
Cited alongside, same era.
J. Ba, R. B. Grosse, and J. Martens, “Distributed second-order optimization using kronecker-factored approximations,” in
2017
Cited alongside, same era.
D. Alistarh, C. De Sa, and N. Konstantinov, “The convergence of stochastic gradient descent in asynchronous shared memory,” in
2018
Later among the works it cites.
J. Martens, J. Ba, and M. Johnson, “Kronecker-factored curvature approximations for recurrent neural networks,” 2018
2018
Later among the works it cites.
T. George, C. Laurent, X. Bouthillier, N. Ballas, and P. Vincent, “Fast approximate natural gradient descent in a kronecker-factored eigenbasis,” in
2018
Later among the works it cites.
K. Osawa, Y. Tsuji, Y. Ueno, A. Naruse, R. Yokota, and S. Matsuoka, “Large-scale distributed second-order optimization using kronecker-factored approximate curvature for deep convolutional neural networks,” in
2019
Later among the works it cites.
H. Lee, M. Turilli, S. Jha, D. Bhowmik, H. Ma, and A. Ramanathan, “DeepDriveMD: Deep-learning driven adaptive molecular simulations for protein folding,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. You, Z. Zhang, C.-J. Hsieh, J. Demmel, and K. Keutzer, “ImageNet training in minutes,” in
2018
Cited alongside, same era.
C. Ying, S. Kumar, D. Chen, T. Wang, and Y. Cheng, “Image classification at supercomputer scale,”
2018
Cited alongside, same era.
H. Mikami, H. Suganuma, Y. Tanaka, Y. Kageyama
2018
Cited alongside, same era.
S. McCandlish, J. Kaplan, D. Amodei, and O. D. Team, “An empirical model of large-batch training,”
2018
Cited alongside, same era.
L. Bottou, F. E. Curtis, and J. Nocedal, “Optimization methods for large-scale machine learning,”
2018
Cited alongside, same era.
2018
Cited alongside, same era.
A. Sergeev and M. D. Balso, “Horovod: Fast and easy distributed deep learning in TensorFlow,”
2018
Cited alongside, same era.
2019
Later among the works it cites.
J. Kates-Harbeck, A. Svyatkovskiy, and W. Tang, “Predicting disruptive instabilities in controlled fusion plasmas through deep learning,”
2019
Later among the works it cites.
Y. You, J. Hseu, C. Ying, J. Demmel, K. Keutzer, and C.-J. Hsieh, “Large-batch training for LSTM and beyond,” in
2019
Later among the works it cites.
L. Ma, G. Montague, J. Ye, Z. Yao, A. Gholami, K. Keutzer, and M. W. Mahoney, “Inefficiency of k-fac for large batch size training,” 2019
2019
Later among the works it cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “Pytorch: An imperative style, high-performance deep learning library,” in
2019
Later among the works it cites.
Intel, “Intel(r) machine learning scaling library,” 2019,
2019
Later among the works it cites.
C. Wang, R. Grosse, S. Fidler, and G. Zhang, “EigenDamage: Structured pruning in the Kronecker-factored eigenbasis,” in
2019
Later among the works it cites.
J. Martens, “Kfac-tensorflow,” 2019. [Online]. Available:
2019
Later among the works it cites.