Fetching the paper…
Reading the bibliography…
Background: Recent developments have made it possible to accelerate neural networks training significantly using large batch sizes and data parallelism.
Gradient-Based Learning Applied to Document Recognition
Yann Lecun, Leon Bottou, Yoshua Bengio, and Patrick Ha · 1998
Earlier work this paper cites.
Discrete-time signal processing
Alan V Oppenheim · 1999
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, and Quoc V Le · 2012
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Scaling Distributed Machine Learning with the Parameter Server
Mu Li · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Asynchronous parallel stochastic gradient for nonconvex optimization
Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu · 2015
Earlier work this paper cites.
Staleness-aware async-sgd for distributed deep learning
Wei Zhang, Suyog Gupta, Xiangru Lian, and Ji Liu · 2015
Earlier work this paper cites.
Revisiting distributed synchronous sgd
Jianmin Chen, Xinghao Pan, Rajat Monga, Samy Bengio, and Rafal Jozefowicz · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
A comprehensive linear speedup analysis for asynchronous stochastic parallel optimization from zeroth-order to first-order
Xiangru Lian, Huan Zhang, Cho-Jui Hsieh, Yijun Huang, and Ji Liu · 2016
Cited alongside, same era.
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
Toward Understanding the Impact of Staleness in Distributed Machine Learning, sep 2018
Wei Dai, Yi Zhou, Nanqing Dong, Hao Zhang, and Eric Xing · 2018
Later among the works it cites.
Slow and Stale Gradients Can Win the Race: Error-Runtime Trade-offs in Distributed SGD
Sanghamitra Dutta, Gauri Joshi, Soumyadip Ghosh, Parijat Dube, and Priya Nagpurkar · 2018
Later among the works it cites.
Towards understanding acceleration tradeoff between momentum and asynchrony in nonconvex stochastic optimization
Tianyi Liu, Shiyang Li, Jianping Shi, Enlu Zhou, and Tuo Zhao · 2018
Later among the works it cites.
Step Size Matters in Deep Learning
Kamil Nar and S Shankar Sastry · 2018
Later among the works it cites.
Measuring the effects of data parallelism on neural network training
Christopher J Shallue, Jaehoon Lee, Joe Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E Dahl · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Accurate, large minibatch sgd: training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
Asynchronous decentralized parallel stochastic gradient descent
Xiangru Lian, Wei Zhang, Ce Zhang, and Ji Liu · 2017
Cited alongside, same era.
Asynchrony begets momentum, with an application to deep learning
Ioannis Mitliagkas, Ce Zhang, Stefan Hadjis, and Christopher Re · 2017
Cited alongside, same era.
A Tight Convergence Analysis for Stochastic Gradient Descent with Delayed Updates
Yossi Arjevani, Ohad Shamir, and Nathan Srebro · 2018
Cited alongside, same era.
How SGD Selects the Global Minima in Over-parameterized Learning: A Dynamical Stability Perspective
Lei Wu, Chao Ma, and Weinan E · 2018
Later among the works it cites.
A unified analysis of stochastic momentum methods for deep learning
Yan Yan, Tianbao Yang, Zhe Li, Qihang Lin, and Yi Yang · 2018
Later among the works it cites.
Ai and compute
Dario Amodei and Danny Hernandez · 2019
Closest in time.
Stochastic Gradient Push for Distributed Deep Learning
Mahmoud Assran, Nicolas Loizou, Nicolas Ballas, and Michael Rabbat · 2019
Closest in time.
Taming Momentum in a Distributed Asynchronous Environment
Ido Hakimi, Saar Barkai, Moshe Gabel, and Assaf Schuster · 2019
Closest in time.