Fetching the paper…
Reading the bibliography…
We demonstrate that training ResNet-50 on ImageNet for 90 epochs can be achieved in 15 minutes with 1024 Tesla P100 GPUs.
Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (ELUs)
D. Clevert, T. Unterthiner, and S. Hochreiter · 2015
Earlier work this paper cites.
Chainer: a next-generation open source framework for deep learning
S. Tokui, K. Oono, S. Hido, and J. Clayton · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, J. Klingner, A. Shah, M. Johnson, X. Liu, L. Kaiser, S. Gouws, Y. Kato, T. Kudo, H. Kazawa, K. Stevens, G. Kurian, N. Patil, W. Wang, C. Young, J. Smith, J. Riesa, A. Rudnick, O. Vinyals, G. Corrado, M. Hughes, and J. Dean · 2016
Cited alongside, same era.
ChainerMN: scalable distributed deep learning framework
T. Akiba, K. Fukuda, and S. Suzuki · 2017
Cited alongside, same era.
Achieving deep learning training in less than 40 minutes on imagenet-1k
V. Codreanu, D. Podareanu, and V. Saletore · 2017
Closest in time.
Accurate, large minibatch SGD: training ImageNet in 1 hour
P. Goyal, P. Dollár, R. B. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Closest in time.
Y. You, Z. Zhang, C. Hsieh, J. Demmel, and K. Keutzer · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…