Fetching the paper…
Reading the bibliography…
Training large deep neural networks on massive datasets is computationally very challenging.
A method for solving the convex programming problem with convergence rate o (1/kˆ 2)
Yurii E Nesterov · 1983
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Practical recommendations for gradient-based training of deep architectures
Yoshua Bengio · 2012
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al · 2012
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
Saeed Ghadimi, Guanghui Lan, and Hongchao Zhang · 2014
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Alex Krizhevsky · 2014
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Incorporating nesterov momentum into adam
Timothy Dozat · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Firecaffe: near-linear acceleration of deep neural network training on compute clusters
Forrest N Iandola, Matthew W Moskewicz, Khalid Ashraf, and Kurt Keutzer · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
Extremely large minibatch sgd: Training resnet-50 on imagenet in 15 minutes
Takuya Akiba, Shuji Suzuki, and Keisuke Fukuda · 2017
Cited alongside, same era.
Valeriu Codreanu, Damian Podareanu, and Vikram Saletore · 2017
Cited alongside, same era.
Dawnbench: An end-to-end deep learning benchmark and competition
Cody Coleman, Deepak Narayanan, Daniel Kang, Tian Zhao, Jian Zhang, Luigi Nardi, Peter Bailis, Kunle Olukotun, Chris Ré, and Matei Zaharia · 2017
Cited alongside, same era.
Adabatch: Adaptive batch sizes for training deep neural networks
signsgd: compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Anima Anandkumar · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Later among the works it cites.
Xianyan Jia, Shutao Song, Wei He, Yangzihao Wang, Haidong Rong, Feihu Zhou, Liqiang Xie, Zhenyu Guo, Yuanzhou Yang, Liwei Yu, et al · 2018
Later among the works it cites.
Imagenet/resnet-50 training in 224 seconds
Hiroaki Mikami, Hisahiro Suganuma, Yoshiki Tanaka, Yuichi Kageyama, et al · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aditya Devarakonda, Maxim Naumov, and Michael Garland · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
Scaling Distributed Machine Learning with System and Algorithm Co-design
Mu Li · 2017
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, and Quoc V Le · 2017
Cited alongside, same era.
Scaling sgd batch size to 32k for imagenet training
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Cited alongside, same era.
Stochastic first- and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan
Cited in the paper.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan
Cited in the paper.
Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno, Akira Naruse, Rio Yokota, and Satoshi Matsuoka · 2018
Later among the works it cites.
Measuring the effects of data parallelism on neural network training
Christopher J Shallue, Jaehoon Lee, Joe Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E Dahl · 2018
Later among the works it cites.
Image classification at supercomputer scale
Chris Ying, Sameer Kumar, Dehao Chen, Tao Wang, and Youlong Cheng · 2018
Later among the works it cites.
Imagenet training in minutes
Yang You, Zhao Zhang, Cho-Jui Hsieh, James Demmel, and Kurt Keutzer · 2018
Later among the works it cites.
Yet another accelerated sgd: Resnet-50 training on imagenet in 74.7 seconds
Masafumi Yamazaki, Akihiko Kasagi, Akihiro Tabuchi, Takumi Honda, Masahiro Miwa, Naoto Fukumoto, Tsuguchika Tabaru, Atsushi Ike, and Kohta Nakashima · 2019
Closest in time.
Large-batch training for lstm and beyond
Yang You, Jonathan Hseu, Chris Ying, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2019
Closest in time.