Fetching the paper…
Reading the bibliography…
We analyze (stochastic) gradient descent (SGD) with delayed updates on smooth quasi-convex and non-convex functions and derive concise, non-asymptotic, convergence rates.
Unified optimal analysis of the (stochastic) gradient method
Sebastian U. Stich · 1907
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Problem Complexity and Method Efficiency in Optimization
A. S. Nemirovski and D.B. Yudin · 1983
Earlier work this paper cites.
Introduction to Optimization
Boris T. Polyak · 1987
Earlier work this paper cites.
Parallel and Distributed Computation: Numerical Methods
D.P. Bertsekas and J.N. Tsitsiklis · 1989
Earlier work this paper cites.
New method of stochastic approximation type
B. T. Polyak · 1990
Earlier work this paper cites.
Introductory Lectures on Convex Optimization , volume 87 of Springer Science & Business Media
Yurii Nesterov · 2004
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Yurii Nesterov and B.T. Polyak · 2006
Earlier work this paper cites.
Efficient large-scale distributed training of conditional maximum entropy models
Ryan McDonald, Mehryar Mohri, Nathan Silberman, Dan Walker, and Gideon S. Mann · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J. Smola · 2010
Earlier work this paper cites.
Distributed delayed stochastic optimization
Alekh Agarwal and John C Duchi · 2011
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Francis R. Bach and Eric Moulines · 2011
Earlier work this paper cites.
HOGWILD!: A lock-free approach to parallelizing stochastic gradient descent
Feng Niu, Benjamin Recht, Christopher Re, and Stephen J. Wright · 2011
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir, and Lin Xiao · 2012
Earlier work this paper cites.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization i: A generic algorithmic framework
Saeed Ghadimi and Guanghui Lan · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Fast convergence of stochastic gradient descent under a strong growth condition
Mark Schmidt and Nicolas Le Roux · 2013
Earlier work this paper cites.
Communication-efficient algorithms for statistical optimization
Yuchen Zhang, John C. Duchi, and Martin J. Wainwright · 2013
Earlier work this paper cites.
Distributed stochastic optimization and learning
O. Shamir and N. Srebro · 2014
Earlier work this paper cites.
Communication complexity of distributed convex learning and optimization
Yossi Arjevani and Ohad Shamir · 2015
Earlier work this paper cites.
Asynchronous stochastic convex optimization: the noise is in the noise and SGD don’t care
Sorathan Chaturapruek, John C Duchi, and Christopher Ré · 2015
Earlier work this paper cites.
INTERSPEECH 2015, 16th Annual Conference of the International Speech Communication Association, Dresden, Germany, September 6-10, 2015 , 2015. ISCA
Haizhou Li, Helen M. Meng, Bin Ma, Engsiong Chng, and Lei Xie, editors · 2015
Cited alongside, same era.
Asynchronous parallel stochastic gradient for nonconvex optimization
Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu · 2015
Cited alongside, same era.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2015
Cited alongside, same era.
Scalable distributed DNN training using commodity GPU cloud computing
Nikko Strom · 2015
Cited alongside, same era.
An asynchronous mini-batch algorithm for regularized stochastic optimization
H. R. Feyzmahdavian, A. Aytekin, and M. Johansson · 2016
Cited alongside, same era.
Parallelizing stochastic gradient descent for least squares regression: Mini-batching, averaging, and model misspecification
Prateek Jain, Sham M. Kakade, Rahul Kidambi, Praneeth Netrapalli, and Aaron Sidford · 2018
Later among the works it cites.
Improved asynchronous parallel optimization analysis for stochastic incremental methods
Remi Leblond, Fabian Pedregosa, and Simon Lacoste-Julien · 2018
Later among the works it cites.
The power of interpolation: Understanding the effectiveness of SGD in modern over-parametrized learning
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2018
Later among the works it cites.
Sparsified SGD with memory
Sebastian U Stich, Jean-Baptiste Cordonnier, and Martin Jaggi · 2018
Later among the works it cites.
Fast and faster convergence of sgd for over-parameterized models and an accelerated perceptron
Sharan Vaswani, Francis Bach, and Mark Schmidt · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Linear convergence of gradient and proximal-gradient methods under the Polyak-Łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Cited alongside, same era.
Optimizing star-convex functions
J. C. H. Lee and P. Valiant · 2016
Cited alongside, same era.
Stochastic gradient descent, weighted sampling, and the randomized Kaczmarz algorithm
Deanna Needell, Nathan Srebro, and Rachel Ward · 2016
Cited alongside, same era.
Adadelay: Delay adaptive distributed stochastic optimization
Suvrit Sra, Adams Wei Yu, Mu Li, and Alex Smola · 2016
Cited alongside, same era.
Parallel SGD: When does averaging help?
Jian Zhang, Christopher De Sa, Ioannis Mitliagkas, and Christopher Ré · 2016
Cited alongside, same era.
QSGD: Communication-efficient SGD via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
Lower bounds for finding stationary points i
Yair Carmon, John C Duchi, Oliver Hinder, and Aaron Sidford · 2017
Cited alongside, same era.
Jianyu Wang and Gauri Joshi · 2018
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang · 2018
Later among the works it cites.
Graph oracle models, lower bounds, and gaps for parallel stochastic optimization
Blake E Woodworth, Jialei Wang, Adam Smith, Brendan McMahan, and Nati Srebro · 2018
Later among the works it cites.
Error compensated quantized SGD and its applications to large-scale distributed optimization
Jiaxiang Wu, Weidong Huang, Junzhou Huang, and Tong Zhang · 2018
Later among the works it cites.
Parallel restarted SGD for non-convex optimization with faster convergence and less communication
Hao Yu, Sen Yang, and Shenghuo Zhu · 2018
Later among the works it cites.
Lower bounds for non-convex stochastic optimization
Yossi Arjevani, Yair Carmon, John C. Duchi, Dylan J. Foster, Nathan Srebro, and Blake Woodworth · 2019
Closest in time.
SGD: General analysis and improved rates
Robert M. Gower, Nicolas Loizou, Xun Qian, Alibek Sailanbayev, Egor Shulgin, and Peter Richtárik · 2019
Closest in time.
Near-optimal methods for minimizing star-convex functions and beyond
Oliver Hinder, Aaron Sidford, and Nimit S. Sohoni · 2019
Closest in time.
Error feedback fixes SignSGD and other gradient compression schemes
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian Stich, and Martin Jaggi · 2019
Closest in time.
Linear convergence of first order methods for non-strongly convex optimization
I. Necoara, Yu. Nesterov, and F. Glineur · 2019
Closest in time.
Communication trade-offs for synchronized distributed SGD with large step size
Kumar Kshitij Patel and Aymeric Dieuleveut · 2019
Closest in time.
A tight convergence analysis for stochastic gradient descent with delayed updates
Yossi Arjevani, Ohad Shamir, and Nathan Srebro · 2020
Closest in time.
SCAFFOLD: stochastic controlled averaging for on-device federated learning
Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi, Sebastian U. Stich, and Ananda Theertha Suresh · 2020
Closest in time.
A unified theory of decentralized SGD with changing topology and local updates
Anastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi, and Sebastian U Stich · 2020
Closest in time.
Don’t use large mini-batches, use local SGD
Tao Lin, Sebastian U. Stich, Kumar Kshitij Patel, and Martin Jaggi · 2020
Closest in time.
On communication compression for distributed optimization on heterogeneous data
Sebastian U. Stich · 2020
Closest in time.
Is local SGD better than minibatch sgd?
Blake E. Woodworth, Kumar Kshitij Patel, Sebastian U. Stich, Zhen Dai, Brian Bullins, H. Brendan McMahan, Ohad Shamir, and Nathan Srebro · 2020
Closest in time.