Fetching the paper…
Reading the bibliography…
We introduce a new and increasingly relevant setting for distributed optimization in machine learning, where the data defining the optimization are unevenly distributed over an extremely large number of nodes.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
Distributed asynchronous computation of fixed points
Dimitri P Bertsekas · 1983
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate
Yurii Nesterov · 1983
Earlier work this paper cites.
Problems in decentralized decision making and computation
John Nikolas Tsitsiklis · 1984
Earlier work this paper cites.
Parallel and distributed computation: numerical methods
Dimitri P Bertsekas and John N Tsitsiklis · 1989
Earlier work this paper cites.
On the limited memory BFGS method for large scale optimization
Dong C Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
An overview of statistical learning theory
Vladimir N Vapnik · 1999
Earlier work this paper cites.
Introductory Lectures on Convex Optimization. A Basic Course
Yurii Nesterov · 2004
Earlier work this paper cites.
The tradeoffs of large scale learning
Olivier Bousquet and Léon Bottou · 2008
Earlier work this paper cites.
MapReduce: Simplified data processing on large clusters
Jeffrey Dean and Sanjay Ghemawat · 2008
Earlier work this paper cites.
Sgd-qn: Careful quasi-Newton stochastic gradient descent
Antoine Bordes, Léon Bottou, and Patrick Gallinari · 2009
Earlier work this paper cites.
Curiously fast convergence of some stochastic gradient descent algorithms
Léon Bottou · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
Spark: cluster computing with working sets
Matei Zaharia, Mosharaf Chowdhury, Michael J Franklin, Scott Shenker, and Ion Stoica · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J Smola · 2010
Earlier work this paper cites.
Distributed delayed stochastic optimization
Alekh Agarwal and John C Duchi · 2011
Earlier work this paper cites.
Scaling up machine learning: Parallel and distributed approaches
Ron Bekkerman, Mikhail Bilenko, and John Langford · 2011
Earlier work this paper cites.
Distributed optimization and statistical learning via the alternating direction method of multipliers
Stephen Boyd, Neal Parikh, Eric Chu, Borja Peleato, and Jonathan Eckstein · 2011
Earlier work this paper cites.
Parallel coordinate descent for l1-regularized loss minimization
Joseph Bradley, Aapo Kyrola, Daniel Bickson, and Carlos Guestrin · 2011
Earlier work this paper cites.
Differentially private empirical risk minimization
Kamalika Chaudhuri, Claire Monteleoni, and Anand D Sarwate · 2011
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Eric Moulines and Francis R Bach · 2011
Earlier work this paper cites.
On optimization methods for deep learning
Jiquan Ngiam, Adam Coates, Ahbik Lahiri, Bobby Prochnow, Quoc V Le, and Andrew Y Ng · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Feng Niu, Benjamin Recht, Christopher Re, and Stephen Wright · 2011
Earlier work this paper cites.
Pegasos: Primal estimated sub-gradient solver for svm
Shai Shalev-Shwartz, Yoram Singer, Nathan Srebro, and Andrew Cotter · 2011
Earlier work this paper cites.
Exascale computing technology challenges
John Shalf, Sudip Dosanjh, and John Morrison · 2011
Earlier work this paper cites.
Stochastic gradient descent tricks
Léon Bottou · 2012
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al · 2012
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir, and Lin Xiao · 2012
Earlier work this paper cites.
Dual averaging for distributed optimization: convergence analysis and network scaling
John C Duchi, Alekh Agarwal, and Martin J Wainwright · 2012
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Efficiency of coordinate descent methods on huge-scale optimization problems
Yu Nesterov · 2012
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence rate for finite training sets
Nicolas Le Roux, Mark Schmidt, and Francis Bach · 2012
Earlier work this paper cites.
Where (and when) do you use your smartphone: Bedroom? Church?
CNN · 2013
Earlier work this paper cites.
Estimation, optimization, and parallelism when data is sparse
John C Duchi, Michael I Jordan, and Brendan H McMahan · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Semi-stochastic gradient descent methods
Jakub Konečný and Peter Richtárik · 2013
Cited alongside, same era.
Efficient accelerated coordinate descent methods and faster algorithms for solving linear systems
Yin Tat Lee and Aaron Sidford · 2013
Cited alongside, same era.
An efficient distributed learning algorithm based on effective local functional approximations
Dhruv Mahajan, Nikunj Agrawal, S Sathiya Keerthi, S Sundararajan, and Leon Bottou · 2013
Cited alongside, same era.
Minimizing finite sums with the stochastic average gradient
Mark Schmidt, Nicolas Le Roux, and Francis Bach · 2013
Cited alongside, same era.
Stochastic dual coordinate ascent methods for regularized loss
Shai Shalev-Shwartz and Tong Zhang · 2013
Cited alongside, same era.
Asynchronous stochastic coordinate descent: Parallelism and convergence properties
Ji Liu and Stephen J Wright · 2015
Later among the works it cites.
An asynchronous parallel stochastic coordinate descent algorithm
Ji Liu, Stephen J Wright, Christopher Ré, Victor Bittorf, and Srikrishna Sridhar · 2015
Later among the works it cites.
Distributed optimization with arbitrary local solvers
Chenxin Ma, Jakub Konečný, Martin Jaggi, Virginia Smith, Michael I Jordan, Peter Richtárik, and Martin Takáč · 2015
Later among the works it cites.
Adding vs. averaging in distributed primal-dual optimization
Chenxin Ma, Virginia Smith, Martin Jaggi, Michael Jordan, Peter Richtárik, and Martin Takáč · 2015
Later among the works it cites.
Perturbed iterate analysis for asynchronous stochastic optimization
Horia Mania, Xinghao Pan, Dimitris Papailiopoulos, Benjamin Recht, Kannan Ramchandran, and Michael I Jordan · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mini-batch primal and dual methods for SVMs
Martin Takáč, Avleen Bijral, Peter Richtárik, and Nathan Srebro · 2013
Cited alongside, same era.
Consumer data privacy in a networked world: A framework for protecting privacy and promoting innovation in the global digital economy
White House Report · 2013
Cited alongside, same era.
Trading computation for communication: Distributed stochastic dual coordinate ascent
Tianbao Yang · 2013
Cited alongside, same era.
Information-theoretic lower bounds for distributed statistical estimation with communication constraints
Yuchen Zhang, John Duchi, Michael I Jordan, and Martin J Wainwright · 2013
Cited alongside, same era.
Communication-efficient algorithms for statistical optimization
Yuchen Zhang, John C Duchi, and Martin J Wainwright · 2013
Cited alongside, same era.
Project adam: Building an efficient and scalable deep learning training system
Trishul Chilimbi, Yutaka Suzue, Johnson Apacible, and Karthik Kalyanaraman · 2014
Cited alongside, same era.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Cited alongside, same era.
Later among the works it cites.
Distributed block coordinate descent for minimizing partially separable functions
Jakub Mareček, Peter Richtárik, and Martin Takáč · 2015
Later among the works it cites.
Arock: an algorithmic framework for asynchronous parallel coordinate updates
Zhimin Peng, Yangyang Xu, Ming Yan, and Wotao Yin · 2015
Later among the works it cites.
Quartz: Randomized dual coordinate ascent with arbitrary sampling
Zheng Qu, Peter Richtárik, and Tong Zhang · 2015
Later among the works it cites.
On variance reduction in stochastic gradient descent and its asynchronous variants
Sashank J Reddi, Ahmed Hefny, Suvrit Sra, Barnabás Póczós, and Alex Smola · 2015
Later among the works it cites.
Non-uniform stochastic average gradient method for training conditional random fields
Mark Schmidt, Reza Babanezhad, Mohamed Ahmed, Aaron Defazio, Ann Clifton, and Anoop Sarkar · 2015
Later among the works it cites.
L1-regularized distributed optimization: A communication-efficient primal-dual framework
Virginia Smith, Simone Forte, Michael I Jordan, and Martin Jaggi · 2015
Later among the works it cites.
Martin Takáč, Peter Richtárik, and Nathan Srebro · 2015
Later among the works it cites.
DiSCO: Distributed optimization for self-concordant empirical loss
Yuchen Zhang and Xiao Lin · 2015
Later among the works it cites.
Distributed newton methods for regularized logistic regression
Yong Zhuang, Wei-Sheng Chin, Yu-Chin Juan, and Chih-Jen Lin · 2015
Later among the works it cites.
Deep learning with differential privacy
Martín Abadi, Andy Chu, Ian Goodfellow, Brendan H McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang · 2016
Closest in time.
Katyusha: The first direct acceleration of stochastic gradient methods
Zeyuan Allen-Zhu · 2016
Closest in time.
Exploiting the structure: Stochastic gradient methods using raw clusters
Zeyuan Allen-Zhu, Yang Yuan, and Karthik Sridharan · 2016
Closest in time.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2016
Closest in time.
A stochastic quasi-Newton method for large-scale optimization
Richard H Byrd, Samantha L Hansen, Jorge Nocedal, and Yoram Singer · 2016
Closest in time.
Coordinate descent face-off: primal or dual?
Dominik Csiba and Peter Richtárik · 2016
Closest in time.
A simple practical accelerated method for finite sums
Aaron Defazio · 2016
Closest in time.
On the global and linear convergence of the generalized alternating direction method of multipliers
Wei Deng and Wotao Yin · 2016
Closest in time.
Stochastic block BFGS: squeezing more curvature out of data
Robert Mansel Gower, Donald Goldfarb, and Peter Richtárik · 2016
Closest in time.
Randomized quasi-Newton updates are linearly convergent matrix inversion algorithms
Robert Mansel Gower and Peter Richtárik · 2016
Closest in time.
Mini-batch semi-stochastic gradient descent in the proximal setting
Jakub Konečný, Jie Liu, Peter Richtárik, and Martin Takáč · 2016
Closest in time.
ASAGA: Asynchronous parallel saga
Rémi Leblond, Fabian Pedregosa, and Simon Lacoste-Julien · 2016
Closest in time.
Federated learning of deep networks using model averaging
Brendan H McMahan, Eider Moore, Daniel Ramage, and Blaise Aguera y Arcas · 2016
Closest in time.
A linearly-convergent stochastic l-bfgs algorithm
Philipp Moritz, Robert Nishihara, and Michael Jordan · 2016
Closest in time.
SDNA: Stochastic dual newton ascent for empirical risk minimization
Zheng Qu, Peter Richtárik, Martin Takáč, and Olivier Fercoq · 2016
Closest in time.
Stochastic variance reduction for nonconvex optimization
Sashank J Reddi, Ahmed Hefny, Suvrit Sra, Barnabás Póczós, and Alex Smola · 2016
Closest in time.
AIDE: Fast and communication efficient distributed optimization
Sashank J Reddi, Jakub Konečný, Peter Richtárik, Barnabás Póczós, and Alex Smola · 2016
Closest in time.
Parallel coordinate descent methods for big data optimization
Peter Richtárik and Martin Takáč · 2016
Closest in time.
Distributed coordinate descent method for learning with big data
Peter Richtárik and Martin Takáč · 2016
Closest in time.
SDCA without duality, regularization, and individual convexity
Shai Shalev-Shwartz · 2016
Closest in time.
Tight complexity bounds for optimizing composite objectives
Blake Woodworth and Nathan Srebro · 2016
Closest in time.