Fetching the paper…
Reading the bibliography…
Distributed optimization is vital in solving large-scale machine learning problems.
Bounds on expectations of linear systematic statistics based on dependent samples
Barry C. Arnold and Richard A. Groeneveld · 1979
Earlier work this paper cites.
Distributed asynchronous deterministic and stochastic gradient optimization algorithms
John Tsitsiklis, Dimitri Bertsekas, and Michael Athans · 1986
Earlier work this paper cites.
Tight bounds on expected order statistics
Dimitris Bertsimas, Karthik Natarajan, and Chung-Piaw Teo · 2006
Earlier work this paper cites.
Distributed subgradient methods for multi-agent optimization
Angelia Nedic and Asuman Ozdaglar · 2009
Earlier work this paper cites.
Primal-dual subgradient methods for convex problems
Yurii Nesterov · 2009
Earlier work this paper cites.
Dual averaging methods for regularized stochastic learning and online optimization
Lin Xiao · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J. Smola · 2010
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg S. Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Quoc V. Le, Mark Z. Mao, Marc’Aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, and Andrew Y. Ng · 2012
Cited alongside, same era.
Optimal distributed online prediction using mini-batches
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir, and Lin Xiao · 2012
Cited alongside, same era.
Dual averaging for distributed optimization: Convergence analysis and network scaling
John C Duchi, Alekh Agarwal, and Martin J Wainwright · 2012
Cited alongside, same era.
Push-sum distributed dual averaging for convex optimization
Konstantinos I Tsianos, Sean Lawlor, and Michael G Rabbat · 2012
Cited alongside, same era.
An asynchronous parallel stochastic coordinate descent algorithm
Ji Liu, Stephen J. Wright, Christopher Ré, Victor Bittorf, and Srikrishna Sridhar · 2015
Cited alongside, same era.
Extra: An exact first-order algorithm for decentralized consensus optimization
Speeding up distributed machine learning using codes
K. Lee, M. Lam, R. Pedarsani, D. Papailiopoulos, and K. Ramchandran · 2017
Later among the works it cites.
Distributed mirror descent for stochastic learning over rate-limited networks
M. Nokleby and W. U. Bajwa · 2017
Later among the works it cites.
Revisiting distributed synchronous sgd
Xinghao Pan, Jianmin Chen, Rajat Monga, Samy Bengio, and Rafal Józefowicz · 2017
Later among the works it cites.
Gradient Coding: Avoiding Stragglers in Distributed Learning
R. Tandon, Q. Lei, A. G. Dimakis, and N. Karampatziakis · 2017
Later among the works it cites.
Dextra: A fast algorithm for optimization over directed graphs
Chenguang Xi and Usman A Khan · 2017
Later among the works it cites.
Polynomial codes: An optimal design for high-dimensional coded matrix multiplication
Q. Yu, M. Maddah-Ali, and S. Avestimehr · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wei Shi, Qing Ling, Gang Wu, and Wotao Yin · 2015
Cited alongside, same era.
Efficient distributed online prediction and stochastic optimization with approximate distributed averaging
Konstantinos I Tsianos and Michael G Rabbat · 2016
Cited alongside, same era.
Later among the works it cites.
Slow and stale gradients can win the race: Error-runtime trade-offs in distributed SGD
S. Ghosh P. Dube S. Dutta, G. Joshi and P. Nagpurkar · 2018
Later among the works it cites.