Fetching the paper…
Reading the bibliography…
Machine learning with big data often involves large optimization models.
Convex Analysis
R. T. Rockafellar · 1970
Earlier work this paper cites.
Monotone operators and the proximal point algorithm
R. T. Rockafellar · 1976
Earlier work this paper cites.
Parallel and Distributed Computation: Numerical Methods
D. P. Bertsekas and J. N. Tsitsiklis · 1989
Earlier work this paper cites.
Fundamentals of Convex Analysis
J.-B. Hiriart-Urruty and C. Lemaréchal · 2001
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Y. Nesterov · 2004
Earlier work this paper cites.
Optimal scaling of a gradient method for distributed resource allocation
L. Xiao and S. P. Boyd · 2006
Earlier work this paper cites.
MapReduce: Simplfied data processing on large clusters
J. Dean and S. Ghemawat · 2008
Earlier work this paper cites.
A fast iterative shrinkage-threshold algorithm for linear inverse problems
A. Beck and M. Teboulle · 2009
Earlier work this paper cites.
Distributed subgradient methods for multi-agent optimization
A. Nedić and A. Ozdaglar · 2009
Earlier work this paper cites.
Distributed delayed stochastic optimization
A. Agarwal and J. C. Duchi · 2011
Earlier work this paper cites.
Distributed optimization and statistical learning via the alternating direction method of multipliers
S. P. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein · 2011
Earlier work this paper cites.
A first-order primal-dual algorithm for convex problems with applications to imaging
A. Chambolle and T. Pock · 2011
Earlier work this paper cites.
LIBSVM data: Classification, regression and multi-label
R.-E. Fan and C.-J. Lin · 2011
Earlier work this paper cites.
OpenMP Application Program Interface, Version 3.1
OpenMP Architecture Review Board · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
B. Recht, C. Re, S. Wright, and F. Niu · 2011
Earlier work this paper cites.
Dual averaging for distributed optimization: convergence analysis and network scaling
J. C. Duchi, A. Agarwal, and M. J. Wainwright · 2012
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence rate for finite training sets
N. Le Roux, M. Schmidt, and F. Bach · 2012
Earlier work this paper cites.
MPI: a message-passing interface standard, Version 3.0
MPI Forum · 2012
Earlier work this paper cites.
Efficiency of coordinate descent methods on huge-scale optimization problems
Y. Nesterov · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Earlier work this paper cites.
Gradient methods for minimizing composite functions
Y. Nesterov · 2013
Earlier work this paper cites.
Stochastic dual coordinate ascent methods for regularized loss minimization
S. Shalev-Shwartz and T. Zhang · 2013
Earlier work this paper cites.
Trading computation for communication: Distributed stochastic dual coordinate ascent
T. Yang · 2013
Cited alongside, same era.
Communication-efficient algorithms for statistical optimization
Y. Zhang, J. C. Duchi, and M. J. Wainwright · 2013
Cited alongside, same era.
Large-scale L-BFGS using MapReduce
W. Chen, Z. Wang, and J. Zhou · 2014
Cited alongside, same era.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
A. Defazio, F. Bach, and S. Lacoste-Julien · 2014
Cited alongside, same era.
Communication-efficient distributed dual coordinate ascent
M. Jaggi, V. Smith, M. Takac, J. Terhorst, S. Krishnan, T. Hofmann, and M. I. Jordan · 2014
Cited alongside, same era.
Scaling distributed machine learning with the parameter server
On variance reduction in stochastic gradient descent and its asynchronous variants
S. J. Reddi, A. Hefny, S. Sra, B. Póczós, and A. J. Smola · 2015
Later among the works it cites.
EXTRA: An exact first-order algorithm for decentralized consensus optimization
W. Shi, Q. Ling, G. Wu, and W. Yin · 2015
Later among the works it cites.
Petuum: A new platform for distributed machine learning on big data
E. P. Xing, Q. Ho, W. Dai, J. K. Kim, J. Wei, S. Lee, X. Zheng, P. Xie, A. Kumar, and Y. Yu · 2015
Later among the works it cites.
Doubly stochastic primal-dual coordinate method for bilinear saddle-point problem
A. W. Yu, Q. Lin, and T. Yang · 2015
Later among the works it cites.
DiSCO: Distributed optimization for self-concordant empirical loss
Y. Zhang and L. Xiao · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su · 2014
Cited alongside, same era.
Large-scale logistic regression and linear support vector machines using Spark
C.-Y. Lin, C.-H. Tsai, C.-P. Lee, and C.-J. Lin · 2014
Cited alongside, same era.
An asynchronous parallel stochastic coordinate descent algorithm
J. Liu, S. J. Wright, C. Ré, V. Bittorf, and S. Sridhar · 2014
Cited alongside, same era.
Distributed stochastic optimization of the regularized risk
S. Matsushima, H. Yun, X. Zhang, and S. V. N. Vishwanathan · 2014
Cited alongside, same era.
Delay-tolerant algorithms for asynchronous distributed online learning
H. B. McMahan and M. J. Streeter · 2014
Cited alongside, same era.
Iteration complexity of randomized block-coordinate descent methods for minimizing a composite function
P. Richtárik and M. Takáč · 2014
Cited alongside, same era.
Communication-efficient distributed optimization using an approximate newton-type method
O. Shamir, N. Srebro, and T. Zhang · 2014
Cited alongside, same era.
A. Aytekin, H. R. Feyzmahdavian, and M. Johansson · 2016
Later among the works it cites.
Stochastic variance reduction methods for saddle-point problems
P. Balamurugan and F. Bach · 2016
Later among the works it cites.
MLlib: Machine learning in Apache Spark
X. Meng, J. Bradley, B. Yavuz, E. Sparks, S. Venkataraman, D. Liu, J. Freeman, D. Tsai, M. Amde, S. Owen, D. Xin, R. Xin, M. J. Franklin, R. Zadeh, M. Zaharia, and A. Talwalkar · 2016
Later among the works it cites.
Achieving geometric convergence for distributed optimizaiton over time-varying graphs
A. Nedić, A. Olshevsky, and W. Shi · 2016
Later among the works it cites.
ARock: An algorithmic framework for asynchronous parallel coordinate updates
Z. Peng, Y. Xu, M. Yan, and W. Yin · 2016
Later among the works it cites.
AIDE: Fast and communication efficient distributed optimization
S. J. Reddi, J. Konečný, P. Richtárik, B. Póczós, and A. Smola · 2016
Later among the works it cites.
Parallel coordinate descent methods for big data optimization
P. Richtárik and M. Takáč · 2016
Later among the works it cites.
A primer on monotone operator methods
E. K. Ryu and S. P. Boyd · 2016
Later among the works it cites.
Adadelay: Delay adaptive distributed stochastic optimization
S. Sra, A. W. Yu, M. Li, and A. J. Smola · 2016
Later among the works it cites.
Apache Spark: A unified engine for big data processing
M. Zaharia, R. Xin, P. Wendell, T. Das, M. Armbrust, A. Dave, X. Meng, J. Rosen, S. Venkataraman, M. J. Franklin, A. Ghodsi, J. Gonzalez, S. Shenker, and I. Stoica · 2016
Later among the works it cites.
Katyusha: The first direct acceleration of stochastic gradient methods
Z. Allen-Zhu · 2017
Closest in time.
More iterations per second, same quality — why asynchronous algorithms may drastically outperform traditional ones
R. Hannah and W. Yin · 2017
Closest in time.
Limited-memory common-directions method for distributed optimization and its application on empirical risk minimization
C.-P. Lee, P.-W. Wang, W. Chen, and C.-J. Lin · 2017
Closest in time.
Distributed optimization with arbitrary local solvers
C. Ma, V. Smith, M. Jaggi, M. I. Jordan, P. Richtárik, and M. Takáč · 2017
Closest in time.
Optimal algorithms for smooth and strongly convex distributed optimization in networks
K. Scaman, F. Bach, S. Bubeck, Y. T. Lee, and L. Massoulié · 2017
Closest in time.
Exploiting strong convexity from data with primal-dual first-order algorithms
J. Wang and L. Xiao · 2017
Closest in time.
Stochastic primal-dual coordinate method for regularized empirical risk minimization
Y. Zhang and L. Xiao · 2017
Closest in time.