Fetching the paper…
Reading the bibliography…
In the last few years, various communication compression techniques have emerged as an indispensable tool helping to alleviate the communication bottleneck in distributed learning.
Error feedback fixes SignSGD and other gradient compression schemes
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian U Stich, and Martin Jaggi · 1901
Earlier work this paper cites.
Stochastic distributed learning with gradient quantization and variance reduction
Samuel Horváth, Dmitry Kovalev, Konstantin Mishchenko, Sebastian Stich, and Peter Richtárik · 1904
Earlier work this paper cites.
Natural compression for distributed deep learning
Samuel Horváth, Chen-Yu Ho, Ľudovít Horváth, Atal Narayan Sahu, Marco Canini, and Peter Richtárik · 1905
Earlier work this paper cites.
Scaffold: Stochastic controlled averaging for on-device federated learning
Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi, Sebastian U. Stich, and Ananda Theertha Suresh · 1910
Earlier work this paper cites.
A First Course in order Statistics
Barry C. Arnold, N. Balakrishnan, and H. N. Nagaraja · 1992
Earlier work this paper cites.
Unified analysis of stochastic gradient methods for composite convex and smooth optimization
Ahmed Khaled, Othmane Sebbouh, Nicolas Loizou, Robert M Gower, and Peter Richtárik · 2006
Earlier work this paper cites.
A fast iterative shrinkage-thresholding algorithm for linear inverse problems
Amir Beck and Marc Teboulle · 2009
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87
Yurii Nesterov · 2013
Earlier work this paper cites.
1-bit stochastic gradient descent and application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Federated learning: strategies for improving communication efficiency
Jakub Konečný, H. Brendan McMahan, Felix Yu, Peter Richtárik, Ananda Theertha Suresh, and Dave Bacon · 2016
Earlier work this paper cites.
Parallel coordinate descent methods for big data optimization
Peter Richtárik and Martin Takáč · 2016
Earlier work this paper cites.
QSGD: Communication-efficient sgd via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Earlier work this paper cites.
Katyusha: The first direct acceleration of stochastic gradient methods
Zeyuan Allen-Zhu · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
H Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Agüera y Arcas · 2017
Earlier work this paper cites.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Cited alongside, same era.
ZipML: Training linear models with end-to-end low precision, and a little bit of deep learning
Hantian Zhang, Jerry Li, Kaan Kara, Dan Alistarh, Ji Liu, and Ce Zhang · 2017
Cited alongside, same era.
The convergence of sparsified gradient methods
Dan Alistarh, Torsten Hoefler, Mikael Johansson, Sarit Khirirat, Nikola Konstantinov, and Cédric Renggli · 2018
Cited alongside, same era.
Convex optimization using sparsified stochastic gradient descent with memory
Jean-Baptiste Cordonnier · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Sebastian U. Stich and Sai Praneeth Karimireddy · 2019
Later among the works it cites.
Sparse gradient compression for distributed SGD
Haobo Sun, Yingxia Shao, Jiawei Jiang, Bin Cui, Kai Lei, Yu Xu, and Jiang Wang · 2019
Later among the works it cites.
Fast and faster convergence of SGD for over-parameterized models and an accelerated perceptron
Sharan Vaswani, Francis Bach, and Mark Schmidt · 2019
Later among the works it cites.
PowerSGD: Practical low-rank gradient compression for distributed optimization
Thijs Vogels, Sai Praneeth Karimireddy, and Martin Jaggi · 2019
Later among the works it cites.
Global momentum compression for sparse communication in distributed SGD
Shen-Yi Zhao, Yinpeng Xie, Hao Gao, and Wu-Jun Li · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hyeontaek Lim, David G Andersen, and Michael Kaminsky · 2018
Cited alongside, same era.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and Bill Dally · 2018
Cited alongside, same era.
Sparsified sgd with memory
Sebastian U Stich, Jean-Baptiste Cordonnier, and Martin Jaggi · 2018
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Cited alongside, same era.
Error compensated quantized SGD and its applications to large-scale distributed optimization
Jiaxiang Wu, Weidong Huang, Junzhou Huang, and Tong Zhang · 2018
Cited alongside, same era.
Advances and open problems in federated learning
Peter Kairouz and et al · 2019
Cited alongside, same era.
Federated learning: challenges, methods, and future directions
Tian Li, Anit Kumar Sahu, Ameet Talwalkar, and Virginia Smith · 2019
Cited alongside, same era.
Accordion: Adaptive gradient communication via critical learning regime identification
Saurabh Agarwal, Hongyi Wang, Kangwook Lee, Shivaram Venkataraman, and Dimitris Papailiopoulos · 2020
Closest in time.
Language models are few-shot learners
Tom B. Brown et al · 2020
Closest in time.
Tighter theory for local SGD on identical and heterogeneous data
Ahmed Khaled, Konstantin Mishchenko, and Peter Richtárik · 2020
Closest in time.
Albert: A lite bert for self-supervised learning of language representations
Zhen-Zhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2020
Closest in time.
A unified analysis of stochastic gradient methods for nonconvex federated optimization
Zhize Li and Peter Richtárik · 2020
Closest in time.
On the Convergence of SGD with Biased Gradients
Ahmad Ajalloeian and Sebastian U. Stich · 2021
Closest in time.
EF21 with Bells & Whistles: Practical Algorithmic Extensions of Modern Error Feedback
Ilyas Fatkhullin, Igor Sokolov, Eduard Gorbunov, Zhize Li, and Peter Richtárik · 2021
Closest in time.
A better alternative to error feedback for communication-efficient distributed learning
Samuel Horváth and Peter Richtárik · 2021
Closest in time.
EF21: A New, Simpler, Theoretically Better, and Practically Faster Error Feedback
Peter Richtárik, Igor Sokolov, and Ilyas Fatkhullin · 2021
Closest in time.
Stochastic sign descent methods: New algorithms and better theory
Mher Safaryan and Peter Richtárik · 2021
Closest in time.