Understand
In \citep{Yangnips13}, the author presented distributed stochastic dual coordinate ascent (DisDCA) algorithms for solving large-scale regularized loss minimization.
- Extraordinary performances have been observed and reported for the well-motivated updates, as referred to the practical updates, compared to the naive updates.
- However, no serious analysis has been provided to understand the updates and therefore the convergence rates.
- In the paper, we bridge the gap by providing a theoretical analysis of the convergence rates of the practical DisDCA algorithm.
Built on
Efficient Large-Scale distributed training of conditional maximum entropy models
Mann, Gideon, McDonald, Ryan, Mohri, Mehryar, Silberman, Nathan, and Walker, Dan · 2009
Earlier work this paper cites.
Parallelized stochastic gradient descent
Zinkevich, Martin, Weimer, Markus, Smola, Alexander J., and Li, Lihong · 2010
Earlier work this paper cites.
Similar
Communication-Efficient algorithms for statistical optimization
Zhang, Yuchen, Duchi, John, and Wainwright, Martin · 2012
Cited alongside, same era.
Then
Trading computation for communication: Distributed stochastic dual coordinate ascent
Yang, Tianbao · 2013
Closest in time.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…