Fetching the paper…
Reading the bibliography…
Many machine learning algorithms minimize a regularized risk, and stochastic optimization is widely used for this task.
On the limited memory BFGS method for large scale optimization
D. C. Liu and J. Nocedal · 1989
Earlier work this paper cites.
Parallel and Distributed Computation: Numerical Methods
D. Bertsekas and J. Tsitsiklis · 1997
Earlier work this paper cites.
Incremental subgradient methods for nondifferentiable optimization
A. Nedić and D. P. Bertsekas · 2001
Earlier work this paper cites.
Learning with Kernels
B. Schölkopf and A. J. Smola · 2002
Earlier work this paper cites.
Convex Optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
Introductory Lectures On Convex Optimization: A Basic Course
Y. Nesterov · 2004
Earlier work this paper cites.
Map-reduce for machine learning on multicore
C.-T. Chu, S. K. Kim, Y.-A. Lin, Y. Yu, G. Bradski, A. Y. Ng, and K. Olukotun · 2006
Earlier work this paper cites.
Introducing the webb spam corpus: Using email spam to identify web spam automatically
S. Webb, J. Caverlee, and C. Pu · 2006
Earlier work this paper cites.
Pegasos: Primal estimated sub-gradient solver for SVM
S. Shalev-Shwartz, Y. Singer, and N. Srebro · 2007
Earlier work this paper cites.
LIBLINEAR: A library for large linear classification
R.-E. Fan, J.-W. Chang, C.-J. Hsieh, X.-R. Wang, and C.-J. Lin · 2008
Earlier work this paper cites.
A dual coordinate descent method for large-scale linear SVM
C. J. Hsieh, K. W. Chang, C. J. Lin, S. S. Keerthi, and S. Sundararajan · 2008
Earlier work this paper cites.
The Elements of Statistical Learning
T. Hastie, R. Tibshirani, and J. Friedman · 2009
Earlier work this paper cites.
Slow learners are fast
J. Langford, A. J. Smola, and M. Zinkevich · 2009
Cited alongside, same era.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Cited alongside, same era.
Parallel inference for latent Dirichlet allocation on graphics processing units
F. Yan, N. Xu, and Y. Qi · 2009
Cited alongside, same era.
Distributed optimization and statistical learning via the alternating direction method of multipliers
S. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2010
Cited alongside, same era.
An architecture for parallel topic models
A. J. Smola and S. Narayanamurthy · 2010
Cited alongside, same era.
Large-scale matrix factorization with distributed stochastic gradient descent
R. Gemulla, E. Nijkamp, P. J. Haas, and Y. Sismanis · 2011
Later among the works it cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
B. Recht, C. Re, S. Wright, and F. Niu · 2011
Later among the works it cites.
More effective distributed ML via a stale synchronous parallel parameter server
Q. Ho, J. Cipar, H. Cui, S. Lee, J. K. Kim, P. B. Gibbons, G. A. Gibson, G. Ganger, and E. P. Xing · 2013
Later among the works it cites.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Later among the works it cites.
Parallel stochastic gradient algorithms for large-scale matrix completion
B. Recht and C. Ré · 2013
Later among the works it cites.
Trading computation for communication: Distributed stochastic dual coordinate ascent
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
COFFIN: a computational framework for linear SVMs
S. Sonnenburg and V. Franc · 2010
Cited alongside, same era.
Bundle methods for regularized risk minimization
C. H. Teo, S. V. N. Vishwanthan, A. J. Smola, and Q. V. Le · 2010
Cited alongside, same era.
Parallelized stochastic gradient descent
M. Zinkevich, A. J. Smola, M. Weimer, and L. Li · 2010
Cited alongside, same era.
The tradeoffs of large-scale learning
L. Bottou and O. Bousquet · 2011
Cited alongside, same era.
Parallel coordinate descent for L1-regularized loss minimization
J. Bradley, A. Kyrola, D. Bickson, and C. Guestrin · 2011
Cited alongside, same era.
T. Yang · 2013
Later among the works it cites.
A reliable effective terascale linear learning system
A. Agarwal, O. Chapelle, M. Dudík, and J. Langford · 2014
Closest in time.
Understanding Machine Learning
S. Shalev-Shwartz and S. Ben-David · 2014
Closest in time.
NOMAD: Non-locking, stOchastic Multi-machine algorithm for Asynchronous and Decentralized matrix completion
H. Yun, H.-F. Yu, C.-J. Hsieh, S. V. N. Vishwanathan, and I. S. Dhillon · 2014
Closest in time.
PASSCoDe: Parallel ASynchronous Stochastic dual Coordinate Descent
C.-J. Hsieh, H.-F. Yu, and I. S. Dhillon · 2015
Closest in time.
DiSCO: Distributed optimization for Self-Concordant empirical loss
Y. Zhang and L. Xiao · 2015
Closest in time.