Fetching the paper…
Reading the bibliography…
Mini-batch stochastic gradient descent (SGD) is state of the art in large scale distributed training.
A Stochastic Approximation Method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Advances in kernel methods
John C. Platt · 1999
Earlier work this paper cites.
Efficient large-scale distributed training of conditional maximum entropy models
Ryan McDonald, Mehryar Mohri, Nathan Silberman, Dan Walker, and Gideon S. Mann · 2009
Earlier work this paper cites.
Distributed training strategies for the structured perceptron
Ryan McDonald, Keith Hall, and Gideon Mann · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J. Smola · 2010
Earlier work this paper cites.
Distributed delayed stochastic optimization
Alekh Agarwal and John C Duchi · 2011
Earlier work this paper cites.
Hogwild!: A lock-free approach to parallelizing stochastic gradient descent
Feng Niu, Benjamin Recht, Christopher Re, and Stephen J. Wright · 2011
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir, and Lin Xiao · 2012
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2012
Earlier work this paper cites.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
Ohad Shamir and Tong Zhang · 2013
Earlier work this paper cites.
Communication-efficient algorithms for statistical optimization
Yuchen Zhang, John C. Duchi, and Martin J. Wainwright · 2013
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Distributed stochastic optimization and learning
O. Shamir and N. Srebro · 2014
Earlier work this paper cites.
Improving deep neural network acoustic models using generalized maxout networks
X. Zhang, J. Trmal, D. Povey, and S. Khudanpur · 2014
Earlier work this paper cites.
Iterative parameter mixing for distributed large-margin training of structured predictors for natural language processing
Greg Coppola · 2015
Earlier work this paper cites.
Deep learning with limited numerical precision
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan · 2015
Earlier work this paper cites.
Asynchronous parallel stochastic gradient for nonconvex optimization
Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu · 2015
Earlier work this paper cites.
Taming the wild: A unified analysis of HOG WILD!-style algorithms
Christopher De Sa, Ce Zhang, Kunle Olukotun, and Christopher Ré · 2015
Cited alongside, same era.
Scalable distributed DNN training using commodity GPU cloud computing
Nikko Strom · 2015
Cited alongside, same era.
Deep learning with elastic averaging SGD
Sixin Zhang, Anna E Choromanska, and Yann LeCun · 2015
Cited alongside, same era.
Stochastic optimization with importance sampling for regularized loss minimization
Peilin Zhao and Tong Zhang · 2015
Cited alongside, same era.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al · 2016
Cited alongside, same era.
On-chip training of recurrent neural networks with limited numerical precision
T. Na, J. H. Ko, J. Kung, and S. Mukhopadhyay · 2017
Later among the works it cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Later among the works it cites.
meProp: Sparsified back propagation for accelerated deep learning with reduced overfitting
Xu Sun, Xuancheng Ren, Shuming Ma, and Houfeng Wang · 2017
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang · 2017
Later among the works it cites.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Avleen S Bijral, Anand D Sarwate, and Nathan Srebro · 2016
Cited alongside, same era.
Revisiting distributed synchronous SGD
Jianmin Chen, Rajat Monga, Samy Bengio, and Rafal Józefowicz · 2016
Cited alongside, same era.
Communication quantization for data-parallel training of deep neural networks
N. Dryden, T. Moon, S. A. Jacobs, and B. V. Essen · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
CNTK: Microsoft’s open-source deep-learning toolkit
Frank Seide and Amit Agarwal · 2016
Cited alongside, same era.
Parallel SGD: When does averaging help?
Jian Zhang, Christopher De Sa, Ioannis Mitliagkas, and Christopher Ré · 2016
Cited alongside, same era.
Sparse communication for distributed gradient descent
Alham Fikri Aji and Kenneth Heafield · 2017
Cited alongside, same era.
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Later among the works it cites.
ZipML: Training linear models with end-to-end low precision, and a little bit of deep learning
Hantian Zhang, Jerry Li, Kaan Kara, Dan Alistarh, Ji Liu, and Ce Zhang · 2017
Later among the works it cites.
Parallelizing stochastic gradient descent for least squares regression: Mini-batching, averaging, and model misspecification
Prateek Jain, Sham M. Kakade, Rahul Kidambi, Praneeth Netrapalli, and Aaron Sidford · 2018
Closest in time.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and Bill Dally · 2018
Closest in time.
The power of interpolation: Understanding the effectiveness of SGD in modern over-parametrized learning
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2018
Closest in time.
Sparsified SGD with memory
Sebastian U Stich, Jean-Baptiste Cordonnier, and Martin Jaggi · 2018
Closest in time.
Jianyu Wang and Gauri Joshi · 2018
Closest in time.
Graph oracle models, lower bounds, and gaps for parallel stochastic optimization
Blake E Woodworth, Jialei Wang, Adam Smith, Brendan McMahan, and Nati Srebro · 2018
Closest in time.
Gradient diversity: a key ingredient for scalable distributed learning
Dong Yin, Ashwin Pananjady, Max Lam, Dimitris Papailiopoulos, Kannan Ramchandran, and Peter Bartlett · 2018
Closest in time.
Parallel restarted SGD for non-convex optimization with faster convergence and less communication
Hao Yu, Sen Yang, and Shenghuo Zhu · 2018
Closest in time.
On the convergence properties of a k-step averaging stochastic gradient descent algorithm for nonconvex optimization
Fan Zhou and Guojing Cong · 2018
Closest in time.
Distributed asynchronous optimization with unbounded delays: How slow can you go?
Zhengyuan Zhou, Panayotis Mertikopoulos, Nicholas Bambos, Peter Glynn, Yinyu Ye, Li-Jia Li, and Li Fei-Fei · 2018
Closest in time.