Fetching the paper…
Reading the bibliography…
We develop several new communication-efficient second-order methods for distributed optimization.
Stochastic distributed learning with gradient quantization and variance reduction
Samuel Horváth, Dmitry Kovalev, Konstantin Mishchenko, Sebastian Stich, and Peter Richtárik · 1904
Earlier work this paper cites.
Natural compression for distributed deep learning
Samuel Horváth, Chen-Yu Ho, Ľudovít Horváth, Atal Narayan Sahu, Marco Canini, and Peter Richtárik · 1905
Earlier work this paper cites.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Quasi-newton methods and their application to function minimisation
Charles G Broyden · 1967
Earlier work this paper cites.
A new approach to variable metric algorithms
Rodger Fletcher · 1970
Earlier work this paper cites.
A family of variable-metric methods derived by variational means
Donald Goldfarb · 1970
Earlier work this paper cites.
Conditioning of quasi-Newton methods for function minimization
David F Shanno · 1970
Earlier work this paper cites.
The modification of Newton’s method for unconstrained optimization by bounding cubic terms
Andreas Griewank · 1981
Earlier work this paper cites.
Cubic regularization of Newton method and its global performance
Yurii Nesterov and Boris T. Polyak · 2006
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
On solving trust-region and other regularised subproblems in optimization
Nicholas IM Gould, Daniel P Robinson, and H Sue Thorne · 2010
Earlier work this paper cites.
Scaling up machine learning: Parallel and distributed approaches
Ron Bekkerman, Mikhail Bilenko, and John Langford · 2011
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Mini-batch primal and dual methods for SVMs
Martin Takáč, Avleen Bijral, Peter Richtárik, and Nathan Srebro · 2013
Earlier work this paper cites.
Introduction to Nonlinear Optimization: Theory, Algorithms, and Applications with MATLAB
Amir Beck · 2014
Earlier work this paper cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data- parallel distributed training of speech DNNs
F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu · 2014
Earlier work this paper cites.
Understanding machine learning: from theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
Communication-efficient distributed optimization using an approximate Newton-type method
Ohad Shamir, Nati Srebro, and Tong Zhang · 2014
Cited alongside, same era.
A proximal stochastic gradient method with progressive variance reduction
Lin Xiao and Tong Zhang · 2014
Cited alongside, same era.
A universal catalyst for first-order optimization
Hongzhou Lin, Julien Mairal, and Zaid Harchaoui · 2015
Cited alongside, same era.
Stochastic gradient descent, weighted sampling, and the randomized Kaczmarz algorithm
Deanna Needell, Nathan Srebro, and Rachel Ward · 2015
Cited alongside, same era.
DiSCO: Distributed optimization for self-concordant empirical loss
Stochastic Newton and cubic Newton methods with simple local linear-quadratic rates
Dmitry Kovalev, Konstanting Mishchenko, and Peter Richtárik · 2019
Later among the works it cites.
Adaptive gradient descent without descent
Yura Malitsky and Konstantin Mishchenko · 2019
Later among the works it cites.
Distributed learning with compressed gradient differences
Konstantin Mishchenko, Eduard Gorbunov, Martin Takáč, and Peter Richtárik · 2019
Later among the works it cites.
New versions of Newton method: step-size choice, convergence domain and under-determined equations
Boris Polyak and Andrey Tremba · 2019
Later among the works it cites.
The error-feedback framework: Better rates for SGD with delayed gradients and compressed communication
S. U. Stich and S. P. Karimireddy · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuchen Zhang and Lin Xiao · 2015
Cited alongside, same era.
Stochastic optimization with importance sampling
Peilin Zhao and Tong Zhang · 2015
Cited alongside, same era.
AIDE: fast and communication efficient distributed optimization
Sashank J. Reddi, Jakub Konečný, Peter Richtárik, Barnabás Póczos, and Alex Smola · 2016
Cited alongside, same era.
QSGD: Communication-efficient SGD via gradient quantization and encoding
D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic · 2017
Cited alongside, same era.
Distributed optimization with arbitrary local solvers
Chenxin Ma, Jakub Konečný, Martin Jaggi, Virginia Smith, Michael I. Jordan, Peter Richtárik, and Martin Takáč · 2017
Cited alongside, same era.
Minimizing finite sums with the stochastic average gradient
M. Schmidt, N. Le Roux, and F. Bach · 2017
Cited alongside, same era.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, and H. Li · 2017
Cited alongside, same era.
DoubleSqueeze: Parallel stochastic gradient descent with double-pass error-compensated compression
H. Tang, X. Lian, T. Zhang, and J. Liu · 2019
Later among the works it cites.
On biased compression for distributed learning
Aleksandr Beznosikov, Samuel Horváth, Peter Richtárik, and Mher Safaryan · 2020
Later among the works it cites.
Communication efficient distributed approximate Newton method
Avishek Ghosh, Raj Kumar Maity, Arya Mazumdar, and Kannan Ramchandran · 2020
Later among the works it cites.
Stochastic subspace cubic Newton method
Filip Hanzely, Nikita Doikov, Yurii Nesterov, and Peter Richtarik · 2020
Later among the works it cites.
Tighter theory for local SGD on identical and heterogeneous data
Ahmed Khaled, Konstantin Mishchenko, and Peter Richtárik · 2020
Later among the works it cites.
Fast linear convergence of randomized BFGS
Dmitry Kovalev, Robert M. Gower, Peter Richtárik, and Alexander Rogozin · 2020
Later among the works it cites.
Acceleration for compressed gradient descent in distributed and federated optimization
Zhize Li, Dmitry Kovalev, Xun Qian, and Peter Richtárik · 2020
Later among the works it cites.
Random reshuffling: Simple analysiswith vast improvements
Konstantin Mishchenko, Ahmed Khaled, and Peter Richtárik · 2020
Later among the works it cites.
New versions of newton method: step-size choice, convergence domain and under-determined equations
Boris Polyak and Andrey Tremba · 2020
Later among the works it cites.
Local SGD converges fast and communicates little
Sebastian U. Stich · 2020
Later among the works it cites.
Distributed adaptive Newton methods with globally superlinear convergence
Jiaqi Zhang, Keyou You, and Tamer Başar · 2020
Later among the works it cites.