Fetching the paper…
Reading the bibliography…
We study the fundamental limits to communication-efficient distributed methods for convex learning and optimization, under different assumptions on the information available to individual machines, and the types of functions considered.
Communication complexity of convex optimization
J. Tsitsiklis and Z.-Q. Luo · 1987
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course
Y. Nesterov · 2004
Earlier work this paper cites.
Smooth minimization of non-smooth functions
Y. Nesterov · 2005
Earlier work this paper cites.
Numerical linear algebra in the streaming model
K. Clarkson and D. Woodruff · 2009
Earlier work this paper cites.
Parallelized stochastic gradient descent
M. Zinkevich, M. Weimer, A. Smola, and L. Li · 2010
Earlier work this paper cites.
A reliable effective terascale linear learning system
A. Agarwal, O. Chapelle, M. Dudík, and J. Langford · 2011
Earlier work this paper cites.
Scaling up machine learning: Parallel and distributed approaches
R. Bekkerman, M. Bilenko, and J. Langford · 2011
Earlier work this paper cites.
Distributed optimization and statistical learning via ADMM
S.P. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein · 2011
Earlier work this paper cites.
Better mini-batch algorithms via accelerated gradient methods
A. Cotter, O. Shamir, N. Srebro, and K. Sridharan · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
B. Recht, C. Ré, S. Wright, and F. Niu · 2011
Cited alongside, same era.
Distributed learning, communication complexity and privacy
M.-F. Balcan, A. Blum, S. Fine, and Y. Mansour · 2012
Cited alongside, same era.
Optimal distributed online prediction using mini-batches
O. Dekel, R. Gilad-Bachrach, O. Shamir, and L. Xiao · 2012
Cited alongside, same era.
Dual averaging for distributed optimization: Convergence analysis and network scaling
J. Duchi, A. Agarwal, and M. Wainwright · 2012
Cited alongside, same era.
Topics in random matrix theory
T. Tao · 2012
Cited alongside, same era.
A parallel SGD method with strong convergence
D. Mahajan, S. Keerthy, S. Sundararajan, and L. Bottou · 2013
Communication-efficient algorithms for statistical optimization
Y. Zhang, J. Duchi, and M. Wainwright · 2013
Later among the works it cites.
Improved distributed principal component analysis
M.-F. Balcan, V. Kanchanapally, Y. Liang, and D. Woodruff · 2014
Later among the works it cites.
Competing with the empirical risk minimizer in a single pass
R. Frostig, R. Ge, S. Kakade, and A. Sidford · 2014
Later among the works it cites.
Communication-efficient distributed dual coordinate ascent
M. Jaggi, V. Smith, M. Takác, J. Terhorst, S. Krishnan, T. Hofmann, and M. Jordan · 2014
Later among the works it cites.
Fundamental limits of online and distributed algorithms for statistical learning and estimation
O. Shamir · 2014
Later among the works it cites.
On distributed stochastic optimization and learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Distributed coordinate descent method for learning with big data
P. Richtárik and M. Takác · 2013
Cited alongside, same era.
Trading computation for communication: Distributed SDCA
T. Yang · 2013
Cited alongside, same era.
Better approximation and faster algorithm using proximal average
Y.-L. Yu · 2013
Cited alongside, same era.
O. Shamir and N. Srebro · 2014
Later among the works it cites.
Communication-efficient distributed optimization using an approximate newton-type method
O. Shamir, N. Srebro, and T. Zhang · 2014
Later among the works it cites.
Distributed stochastic variance reduced gradient methods
J. Lee, T. Ma, and Q. Lin · 2015
Closest in time.
Communication-efficient distributed optimization of self-concordant empirical loss
Y. Zhang and L. Xiao · 2015
Closest in time.