Fetching the paper…
Reading the bibliography…
We consider the setting of distributed empirical risk minimization where multiple machines compute the gradients in parallel and a centralized server updates the model parameters.
Probability inequalities for sums of bounded random variables
Wassily Hoeffding · 1963
Earlier work this paper cites.
Convex Analysis
R. T. Rockafellar · 1970
Earlier work this paper cites.
Convergence analysis of a proximal-like minimization algorithm using Bregman functions
Gong Chen and Marc Teboulle · 1993
Earlier work this paper cites.
RCV1: A new benchmark collection for text categorization research
David D Lewis, Yiming Yang, Tony G Rose, and Fan Li · 2004
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Yurii Nesterov · 2004
Earlier work this paper cites.
A fast iterative shrinkage-thresholding algorithm for linear inverse problems
Amir Beck and Marc Teboulle · 2009
Earlier work this paper cites.
Distributed optimization and statistical learning via the alternating direction method of multipliers
Stephen Boyd, Neal Parikh, Eric Chu, Borja Peleato, and Jonathan Eckstein · 2010
Earlier work this paper cites.
Feature engineering and classifier ensemble for KDD cup 2010
Hsiang-Fu Yu, Hung-Yi Lo, Hsun-Ping Hsieh, Jing-Kai Lou, Todd G McKenzie, Jung-Wei Chou, Po-Han Chung, Chia-Hua Ho, Chun-Fu Chang, Yin-Hsuan Wei, et al · 2010
Earlier work this paper cites.
Concentration Inequalities: A Nonasymptotic Theory of Independence
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Gradient methods for minimizing composite functions
Yurii Nesterov · 2013
Earlier work this paper cites.
Communication-efficient distributed dual coordinate ascent
Martin Jaggi, Virginia Smith, Martin Takac, Jonathan Terhorst, Sanjay Krishnan, Thomas Hofmann, and Michael I. Jordan · 2014
Earlier work this paper cites.
Communication-efficient distributed optimization using an approximate Newton-type method
Ohad Shamir, Nati Srebro, and Tong Zhang · 2014
Cited alongside, same era.
Communication complexity of distributed convex learning and optimization
Yossi Arjevani and Ohad Shamir · 2015
Cited alongside, same era.
A universal catalyst for first-order optimization
Hongzhou Lin, Julien Mairal, and Zaid Harchaoui · 2015
Cited alongside, same era.
An adaptive accelerated proximal gradient method and its homotopy continuation for sparse optimization
Qihang Lin and Lin Xiao · 2015
Cited alongside, same era.
Adding vs. averaging in distributed primal-dual optimization
Chenxin Ma, Virginia Smith, Martin Jaggi, Michael I. Jordan, Peter Richtárik, and Martin Takáč · 2015
Cited alongside, same era.
An introduction to matrix concentration inequalities
Joel A. Tropp · 2015
Efficiency of the accelerated coordinate descent method on structured optimization problems
Yurii Nesterov and Sebastian U Stich · 2017
Later among the works it cites.
Optimal algorithms for smooth and strongly convex distributed optimization in networks
Kevin Scaman, Francis Bach, Sébastien Bubeck, Yin Tat Lee, and Laurent Massoulié · 2017
Later among the works it cites.
Accelerated Bregman proximal gradient methods for relatively smooth convex optimization
Filip Hanzely, Peter Richtarik, and Lin Xiao · 2018
Later among the works it cites.
Relatively smooth convex optimization by first-order methods, and applications
Haihao Lu, Robert M Freund, and Yurii Nesterov · 2018
Later among the works it cites.
An efficient distributed learning algorithm based on effective local functional approximations
Dhruv Mahajan, Nikunj Agrawal, S. Sathiya Keerthi, Sundararajan Sellamanickam, and Leon Bottou · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
DiSCO: Distributed optimization for self-concordant empirical loss
Yuchen Zhang and Lin Xiao · 2015
Cited alongside, same era.
AIDE: Fast and communication efficient distributed optimization
Sashank J. Reddi, Jakub Konečnỳ, Peter Richtárik, Barnabás Póczós, and Alex Smola · 2016
Cited alongside, same era.
Sdca without duality, regularization, and individual convexity
Shai Shalev-Shwartz · 2016
Cited alongside, same era.
A descent lemma beyond Lipschitz gradient continuity: first-order methods revisited and applications
Heinz H. Bauschke, Jérôme Bolte, and Marc Teboulle · 2017
Cited alongside, same era.
First-Order Methods in Optimization
Amir Beck · 2017
Cited alongside, same era.
Distributed optimization with arbitrary local solvers
Chenxin Ma, Virginia Smith, Martin Jaggi, Michael I. Jordan, Peter Richtárik, and Martin Takáč · 2017
Cited alongside, same era.
GIANT: Globally improved approximate Newton method for distributed optimization
Shusen Wang, Farbod Roosta-Khorasani, Peng Xu, and Michael W Mahoney · 2018
Later among the works it cites.
Communication-efficient distributed optimization of self-concordant empirical loss
Yuchen Zhang and Lin Xiao · 2018
Later among the works it cites.
Optimal complexity and certification of Bregman first-order methods
Radu-Alexandru Dragomir, Adrien Taylor, Alexandre d’Aspremont, and Jérôme Bolte · 2019
Later among the works it cites.
High-Dimensional Probability, An Introduction with Applications in Data Science
Roman Vershynin · 2019
Later among the works it cites.
Utilizing second order information in minibatch stochastic variance reduced proximal iterations
Jialei Wang and Tong Zhang · 2019
Later among the works it cites.
DSCOVR: Randomized primal-dual block coordinate algorithms for asynchronous distributed optimization
Lin Xiao, Adams Wei Yu, Qihang Lin, and Weizhu Chen · 2019
Later among the works it cites.
On convergence of distributed approximate Newton methods: Globalization, sharper bounds and beyond
Xiao-Tong Yuan and Ping Li · 2019
Later among the works it cites.