Fetching the paper…
Reading the bibliography…
We study distributed stochastic convex optimization under the delayed gradient model where the server nodes perform parameter updates, while the worker nodes compute stochastic gradients.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Parallel and Distributed Computation: Numerical Methods
D. Bertsekas and J. Tsitsiklis · 1989
Earlier work this paper cites.
Distributed asynchronous incremental subgradient methods
A Nedić, Dimitri P Bertsekas, and Vivek S Borkar · 2001
Earlier work this paper cites.
J. Langford, A. J. Smola, and M. Zinkevich · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Earlier work this paper cites.
Distributed stochastic subgradient projection algorithms for convex optimization
S Sundhar Ram, A Nedić, and Venugopal V Veeravalli · 2010
Earlier work this paper cites.
Stochastic Optimization for Machine Learning
N. Srebro and A. Tewari · 2010
Cited alongside, same era.
Distributed delayed stochastic optimization
Alekh Agarwal and John C Duchi · 2011
Cited alongside, same era.
Incremental gradient, subgradient, and proximal methods for convex optimization: A survey
Dimitri P Bertsekas · 2011
Cited alongside, same era.
Dual averaging for distributed optimization: convergence analysis and network scaling
John C Duchi, Alekh Agarwal, and Martin J Wainwright · 2012
Cited alongside, same era.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization i: A generic algorithmic framework
Saeed Ghadimi and Guanghui Lan · 2012
Cited alongside, same era.
Scaling distributed machine learning with the parameter server
Mu Li, David G Andersen, Jun Woo Park, Alexander J Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J Shekita, and Bor-Yiing Su
Estimation, optimization, and parallelism when data is sparse
John Duchi, Michael I Jordan, and Brendan McMahan · 2013
Later among the works it cites.
Minimizing finite sums with the stochastic average gradient
Mark W. Schmidt, Nicolas Le Roux, and Francis R. Bach · 2013
Later among the works it cites.
Delay-tolerant algorithms for asynchronous distributed online learning
Brendan McMahan and Matthew Streeter · 2014
Later among the works it cites.
Distributed stochastic optimization and learning
Ohad Shamir and Nathan Srebro · 2014
Later among the works it cites.
Lectures on stochastic programming: modeling and theory , volume 16
Alexander Shapiro, Darinka Dentcheva, and Andrzej Ruszczyński · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited in the paper.
Communication efficient distributed machine learning with the parameter server
Mu Li, David G Andersen, Alex J Smola, and Kai Yu
Cited in the paper.