Fetching the paper…
Reading the bibliography…
The scale of modern datasets necessitates the development of efficient distributed optimization methods for machine learning.
Parallel and Distributed Computation: Numerical Methods
D. P. Bersekas and J. N. Tsitsiklis · 1989
Earlier work this paper cites.
Convex Analysis
R. T. Rockafellar · 1997
Earlier work this paper cites.
Fundamentals of convex analysis
J.-B. Hiriart-Urruty and C. Lemaréchal · 2001
Earlier work this paper cites.
Convex Optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
Techniques of Variational Analysis
J. M. Borwein and Q. Zhu · 2005
Earlier work this paper cites.
Smooth minimization of non-smooth functions
Y. Nesterov · 2005
Earlier work this paper cites.
Scalable training of L1-regularized log-linear models
G. Andrew and J. Gao · 2007
Earlier work this paper cites.
LIBLINEAR: A library for large linear classification
R.-E. Fan, K.-W. Chang, C.-J. Hsieh, X.-R. Wang, and C.-J. Lin · 2008
Earlier work this paper cites.
Efficient large-scale distributed training of conditional maximum entropy models
G. Mann, R. McDonald, M. Mohri, N. Silberman, and D. D. Walker · 2009
Earlier work this paper cites.
Distributed optimization and statistical learning via the alternating direction method of multipliers
S. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein · 2010
Earlier work this paper cites.
Regularization paths for generalized linear models via coordinate descent
J. Friedman, T. Hastie, and R. Tibshirani · 2010
Earlier work this paper cites.
A quasi-Newton approach to nonsmooth convex optimization problems in machine learning
J. Yu, S. Vishwanathan, S. Günter, and N. N. Schraudolph · 2010
Earlier work this paper cites.
A comparison of optimization methods and software for large-scale l1-regularized linear classification
G.-X. Yuan, K.-W. Chang, C.-J. Hsieh, and C.-J. Lin · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
M. A. Zinkevich, M. Weimer, A. J. Smola, and L. Li · 2010
Earlier work this paper cites.
Convex Analysis and Monotone Operator Theory in Hilbert Spaces
H. H. Bauschke and P. L. Combettes · 2011
Earlier work this paper cites.
Parallel coordinate descent for l1-regularized loss minimization
J. K. Bradley, A. Kyrola, D. Bickson, and C. Guestrin · 2011
Earlier work this paper cites.
Hogwild!: A lock-free approach to parallelizing stochastic gradient descent
F. Niu, B. Recht, C. Ré, and S. J. Wright · 2011
Earlier work this paper cites.
Solving large scale linear SVM with distributed block minimization
D. Pechyony, L. Shen, and R. Jones · 2011
Earlier work this paper cites.
Stochastic methods for l 1
S. Shalev-Shwartz and A. Tewari · 2011
Earlier work this paper cites.
Distributed learning, communication complexity and privacy
M.-F. Balcan, A. Blum, S. Fine, and Y. Mansour · 2012
Earlier work this paper cites.
Optimal Distributed Online Prediction Using Mini-Batches
O. Dekel, R. Gilad-Bachrach, O. Shamir, and L. Xiao · 2012
Earlier work this paper cites.
Large linear classification when data cannot fit in memory
H.-F. Yu, C.-J. Hsieh, K.-W. Chang, and C.-J. Lin · 2012
Earlier work this paper cites.
An improved GLMNET for L1-regularized logistic regression
G.-X. Yuan, C.-H. Ho, and C.-J. Lin · 2012
Earlier work this paper cites.
Efficient distributed linear classification algorithms via the alternating direction method of multipliers
C. Zhang, H. Lee, and K. G. Shin · 2012
Cited alongside, same era.
Parallel coordinate descent Newton method for efficient ℓ \ell 1
Y. Bian, X. Li, Y. Liu, and M.-H. Yang · 2013
Cited alongside, same era.
Estimation, optimization, and parallelism when data is sparse
J. Duchi, M. I. Jordan, and B. McMahan · 2013
Cited alongside, same era.
On the complexity analysis of randomized block-coordinate descent methods
Z. Lu and L. Xiao · 2013
Cited alongside, same era.
D-ADMM: A Communication-Efficient Distributed Algorithm for Separable Optimization
J. F. C. Mota, J. M. F. Xavier, P. M. Q. Aguiar, and M. Puschel · 2013
Cited alongside, same era.
Mini-batch primal and dual methods for SVMs
M. Takáč, A. Bijral, P. Richtárik, and N. Srebro · 2013
Quartz: Randomized dual coordinate ascent with arbitrary sampling
Z. Qu, P. Richtárik, and T. Zhang · 2015
Later among the works it cites.
L1-Regularized Distributed Optimization: A Communication-Efficient Primal-Dual Framework
V. Smith, S. Forte, M. I. Jordan, and M. Jaggi · 2015
Later among the works it cites.
On the complexity of parallel coordinate descent
R. Tappenden, M. Takáč, and P. Richtárik · 2015
Later among the works it cites.
Coordinate descent algorithms
S. J. Wright · 2015
Later among the works it cites.
A dual augmented block minimization framework for learning with limited memory
I. E.-H. Yen, S.-W. Lin, and S.-D. Lin · 2015
Later among the works it cites.
Stochastic primal-dual coordinate method for regularized empirical risk minimization
Y. Zhang and X. Lin · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Trading computation for communication: Distributed stochastic dual coordinate ascent
T. Yang · 2013
Cited alongside, same era.
Analysis of distributed stochastic dual coordinate ascent
T. Yang, S. Zhu, R. Jin, and Y. Lin · 2013
Cited alongside, same era.
Communication-efficient algorithms for statistical optimization
Y. Zhang, J. C. Duchi, and M. J. Wainwright · 2013
Cited alongside, same era.
Communication-efficient distributed dual coordinate ascent
M. Jaggi, V. Smith, M. Takáč, J. Terhorst, S. Krishnan, T. Hofmann, and M. I. Jordan · 2014
Cited alongside, same era.
LOCO: Distributing ridge regression with random projections
B. McWilliams, C. Heinze, N. Meinshausen, G. Krummenacher, and H. P. Vanchinathan · 2014
Cited alongside, same era.
Distributed dual gradient methods and error bound conditions
I. Necoara and V. Nedelcu · 2014
Cited alongside, same era.
Primal-dual rates and certificates
C. Dünner, S. Forte, M. Takáč, and M. Jaggi · 2016
Closest in time.
DUAL-LOCO: Distributing statistical estimation using random projections
C. Heinze, B. McWilliams, and N. Meinshausen · 2016
Closest in time.
Linear convergence of gradient and proximal-gradient methods under the Polyak-łojasiewicz condition
H. Karimi, J. Nutini, and M. Schmidt · 2016
Closest in time.
MLlib: Machine learning in apache spark
X. Meng, J. Bradley, B. Yavuz, E. Sparks, S. Venkataraman, D. Liu, J. Freeman, D. Tsai, M. Amde, S. Owen, D. Xin, R. Xin, M. J. Franklin, R. Zadeh, M. Zaharia, and A. Talwalkar · 2016
Closest in time.
SDNA: Stochastic dual Newton ascent for empirical risk minimization
Z. Qu, P. Richtárik, M. Takáč, and O. Fercoq · 2016
Closest in time.
Distributed coordinate descent method for learning with big data
P. Richtárik and M. Takáč · 2016
Closest in time.
Distributed Coordinate Descent for Generalized Linear Models with Regularization
I. Trofimov and A. Genkin · 2016
Closest in time.
Adabatch: Efficient gradient aggregation rules for sequential and parallel stochastic gradient methods
A. Défossez and F. Bach · 2017
Closest in time.
Efficient use of limited-memory accelerators for linear learning on heterogeneous systems
C. Dünner, T. Parnell, and M. Jaggi · 2017
Closest in time.
Hessian-CoCoA: a general parallel and distributed framework for non-strongly convex regularizers
M. Gargiani · 2017
Closest in time.
Distributed block-diagonal approximation methods for regularized empirical risk minimization
C.-p. Lee and K.-W. Chang · 2017
Closest in time.
A distributed block coordinate descent method for training l 1 regularized linear classifiers
D. Mahajan, S. S. Keerthi, and S. Sundararajan · 2017
Closest in time.
Federated multi-task learning
V. Smith, C.-K. Chiang, M. Sanjabi, and A. S. Talwalkar · 2017
Closest in time.
A general distributed dual coordinate optimization framework for regularized loss minimization
S. Zheng, J. Wang, F. Xia, W. Xu, and T. Zhang · 2017
Closest in time.
A Distributed Second-Order Algorithm You Can Trust
C. Dünner, A. Lucchi, M. Gargiani, A. Bian, T. Hofmann, and M. Jaggi · 2018
Closest in time.
A distributed quasi-newton algorithm for empirical risk minimization with nonsmooth regularization
C.-p. Lee, C. H. Lim, and S. J. Wright · 2018
Closest in time.