Fetching the paper…
Reading the bibliography…
Error feedback (EF), also known as error compensation, is an immensely popular convergence stabilization mechanism in the context of distributed training of supervised machine learning models enhanced by the use of contractive communication compression mechanisms, such as Top-$k$.
Stochastic distributed learning with gradient quantization and variance reduction
Samuel Horváth, Dmitry Kovalev, Konstantin Mishchenko, Sebastian Stich, and Peter Richtárik · 1904
Earlier work this paper cites.
Natural compression for distributed deep learning
Samuel Horváth, Chen-Yu Ho, Ľudovít Horváth, Atal Narayan Sahu, Marco Canini, and Peter Richtárik · 1905
Earlier work this paper cites.
Introductory lectures on convex optimization: a basic course (Applied Optimization)
Yurii Nesterov · 2004
Earlier work this paper cites.
Page: A simple and optimal probabilistic gradient estimator for nonconvex optimization
Zhize Li, Hongyan Bao, Xiangliang Zhang, and Peter Richtárik · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Libsvm: a library for support vector machines
Chih-Chung Chang and Chih-Jen Lin · 2011
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
QSGD: Communication-efficient SGD via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Earlier work this paper cites.
Non-convex optimization for machine learning
Prateek Jain and Purushottam Kar · 2017
Earlier work this paper cites.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Xiangru Lian, Ce Zhang, Huan Zhang, Cho-Jui Hsieh, Wei Zhang, and Ji Liu · 2017
Earlier work this paper cites.
The convergence of sparsified gradient methods
Dan Alistarh, Torsten Hoefler, Mikael Johansson, Sarit Khirirat, Nikola Konstantinov, and Cédric Renggli · 2018
Earlier work this paper cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Cited alongside, same era.
Distributed learning with compressed gradients
Sarit Khirirat, Hamid Reza Feyzmahdavian, and Mikael Johansson · 2018
Cited alongside, same era.
Sparsified SGD with memory
Sebastian U. Stich, J.-B. Cordonnier, and Martin Jaggi · 2018
Cited alongside, same era.
Qsparse-local-SGD: Distributed SGD with quantization, sparsification, and local computations
Debraj Basu, Deepesh Data, Can Karakus, and Suhas Diggavi · 2019
Cited alongside, same era.
Error feedback fixes SignSGD and other gradient compression schemes
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian Stich, and Martin Jaggi · 2019
Cited alongside, same era.
Decentralized deep learning with arbitrary communication compression
Anastasia Koloskova, Tao Lin, S. Stich, and Martin Jaggi · 2020
Later among the works it cites.
A unified analysis of stochastic gradient methods for nonconvex federated optimization
Zhize Li and Peter Richtárik · 2020
Later among the works it cites.
Acceleration for compressed gradient descent in distributed and federated optimization
Zhize Li, Dmitry Kovalev, Xun Qian, and Peter Richtárik · 2020
Later among the works it cites.
99% of worker-master communication in distributed optimization is not needed
Konstantin Mishchenko, Filip Hanzely, and Peter Richtárik · 2020
Later among the works it cites.
Constantin Philippenko and Aymeric Dieuleveut · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Konstantin Mishchenko, Eduard Gorbunov, Martin Takáč, and Peter Richtárik · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Cited alongside, same era.
Sebastian Stich and Sai Praneeth Karimireddy · 2019
Cited alongside, same era.
A survey on distributed machine learning
Joost Verbraeken, Matthijs Wolting, Jonathan Katzy, Jeroen Kloppenburg, Tim Verbelen, and Jan S Rellermeyer · 2019
Cited alongside, same era.
Analysis of SGD with biased gradient estimators
Ahmad Ajalloeian and Sebastian U Stich · 2020
Cited alongside, same era.
On biased compression for distributed learning
Aleksandr Beznosikov, Samuel Horváth, Peter Richtárik, and Mher Safaryan · 2020
Cited alongside, same era.
Unified analysis of stochastic gradient methods for composite convex and smooth optimization
Ahmed Khaled, Othmane Sebbouh, Nicolas Loizou, Robert M. Gower, and Peter Richtárik · 2020
Cited alongside, same era.
Later among the works it cites.
Error compensated distributed SGD can be accelerated
Xun Qian, Peter Richtárik, and Tong Zhang · 2020
Later among the works it cites.
DoubleSqueeze: Parallel stochastic gradient descent with double-pass error-compensated compression
Hanlin Tang, Xiangru Lian, Chen Yu, Tong Zhang, and Ji Liu · 2020
Later among the works it cites.
CSER: Communication-efficient SGD with error reset
Cong Xie, Shuai Zheng, Oluwasanmi Koyejo, Indranil Gupta, Mu Li, and Haibin Lin · 2020
Later among the works it cites.
Compressed communication for distributed deep learning: Survey and quantitative evaluation
Hang Xu, Chen-Yu Ho, Ahmed M Abdelmoniem, Aritra Dutta, El Houcine Bergou, Konstantinos Karatsenidis, Marco Canini, and Panos Kalnis · 2020
Later among the works it cites.
MARINA: Faster non-convex distributed learning with compression
Eduard Gorbunov, Konstantin Burlachenko, Zhize Li, and Peter Richtárik · 2021
Closest in time.
Distributed second order methods with fast rates and compressed communication
Rustem Islamov, Xun Qian, and Peter Richtárik · 2021
Closest in time.
Uncertainty principle for communication compression in distributed and federated learning and the search for an optimal compressor
Mher Safaryan, Egor Shulgin, and Peter Richtárik · 2021
Closest in time.