Fetching the paper…
Reading the bibliography…
Inspired by recent work of Islamov et al (2021), we propose a family of Federated Newton Learn (FedNL) methods, which we believe is a marked step in the direction of making second-order methods applicable to FL.
LibSVM: a library for support vector machines
Chih-Chung Chang and Chih-Jen Lin · 2011
Earlier work this paper cites.
Introduction to Nonlinear Optimization: Theory, Algorithms, and Applications with MATLAB
Amir Beck · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Communication-effcient distributed optimization using an approximate newton-type method
Ohad Shamir, Nati Srebro, and Tong Zhang · 2014
Earlier work this paper cites.
Disco: Distributed optimization for self-concordant empirical loss
Yuchen Zhang and Xiao Lin · 2015
Earlier work this paper cites.
AIDE: Fast and communication efficient distributed optimization
Sashank J. Reddi, Jakub Konečný, Peter Richtárik, Barnabás Póczos, and Alexander J. Smola · 2016
Earlier work this paper cites.
QSGD: Communication-efficient SGD via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
H Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Agüera y Arcas · 2017
Earlier work this paper cites.
Distributed learning with compressed gradients
Sarit Khirirat, Hamid Reza Feyzmahdavian, and Mikael Johansson · 2018
Earlier work this paper cites.
Federated optimization in heterogeneous networks
Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi, Ameet Talwalkar, and Virginia Smith · 2018
Earlier work this paper cites.
Sparsified SGD with memory
S. U. Stich, J.-B. Cordonnier, and M. Jaggi · 2018
Earlier work this paper cites.
GIANT: Globally improved approximate Newton method for distributed optimization
Shusen Wang, Fred Roosta abd Peng Xu, and Michael W Mahoney · 2018
Earlier work this paper cites.
Dingo: Distributed newton-type method for gradient-norm optimization
Rixon Crane and Fred Roosta · 2019
Cited alongside, same era.
SGD: General analysis and improved rates
Robert Mansel Gower, Nicolas Loizou, Xun Qian, Alibek Sailanbayev, Egor Shulgin, and Peter Richtárik · 2019
Cited alongside, same era.
Stochastic distributed learning with gradient quantization and variance reduction
Samuel Horváth, Dmitry Kovalev, Konstantin Mishchenko, Sebastian Stich, and Peter Richtárik · 2019
Cited alongside, same era.
Advances and open problems in federated learning
Peter Kairouz et al · 2019
Cited alongside, same era.
Error feedback fixes SignSGD and other gradient compression schemes
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian Stich, and Martin Jaggi · 2019
Cited alongside, same era.
Linearly converging error compensated SGD
Eduard Gorbunov, Dmitry Kovalev, Dmitry Makarenko, and Peter Richtárik · 2020
Later among the works it cites.
SCAFFOLD: Stochastic controlled averaging for on-device federated learning
Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi, Sebastian U. Stich, and Ananda Theertha Suresh · 2020
Later among the works it cites.
Tighter theory for local SGD on identical and heterogeneous data
Ahmed Khaled, Konstantin Mishchenko, and Peter Richtárik · 2020
Later among the works it cites.
Federated learning: challenges, methods, and future directions
Tian Li, Anit Kumar Sahu, Ameet Talwalkar, and Virginia Smith · 2020
Later among the works it cites.
A double residual compression algorithm for efficient distributed learning
Xiaorui Liu, Yao Li, Jiliang Tang, and Ming Yan · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adaptive gradient descent without descent
Yura Malitsky and Konstantin Mishchenko · 2019
Cited alongside, same era.
Distributed learning with compressed gradient differences
Konstantin Mishchenko, Eduard Gorbunov, Martin Takáč, and Peter Richtárik · 2019
Cited alongside, same era.
PowerSGD: Practical low-rank gradient compression for distributed optimization
Thijs Vogels, Sai Praneeth Karimireddy, and Martin Jaggi · 2019
Cited alongside, same era.
Local AdaAlter: Communication-efficient stochastic gradient descent with adaptive learning rates
Cong Xie, Oluwasanmi Koyejo, Indranil Gupta, and Haibin Lin · 2019
Cited alongside, same era.
On biased compression for distributed learning
Aleksandr Beznosikov, Samuel Horváth, Peter Richtárik, and Mher Safaryan · 2020
Cited alongside, same era.
Optimal client sampling for federated learning
Wenlin Chen, Samuel Horváth, and Peter Richtárik · 2020
Cited alongside, same era.
A unified theory of SGD: Variance reduction, sampling, quantization and coordinate descent
Eduard Gorbunov, Filip Hanzely, and Peter Richtárik
Cited in the paper.
Sashank Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett, Keith Rush, Jakub Konečný, Sanjiv Kumar, and H. Brendan McMahan · 2020
Later among the works it cites.
Local SGD converges fast and communicates little
Sebastian U. Stich · 2020
Later among the works it cites.
Achieving globally superlinear convergence for distributed optimization with adaptive newton method
Jiaqi Zhang, Keyou You, and Tamer Başar · 2020
Later among the works it cites.
Communication-efficient distributed optimization with quantized preconditioners
Foivos Alimisis, Peter Davies, and Dan Alistarh · 2021
Closest in time.
Distributed second order methods with fast rates and compressed communication
Rustem Islamov, Xun Qian, and Peter Richtárik · 2021
Closest in time.
Constantin Philippenko and Aymeric Dieuleveut · 2021
Closest in time.