Fetching the paper…
Reading the bibliography…
In this paper, we propose and analyze SQuARM-SGD, a communication-efficient algorithm for decentralized training of large-scale machine learning models over a network.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Cifar-10
Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton · 2009
Earlier work this paper cites.
Distributed event-triggered control for multi-agent systems
Dimos V. Dimarogonas, Emilio Frazzoli, and Karl Henrik Johansson · 2012
Earlier work this paper cites.
An introduction to event-triggered and self-triggered control
W. P. M. H. Heemels, Karl Henrik Johansson, and Paulo Tabuada · 2012
Earlier work this paper cites.
Event-based broadcasting for multi-agent average consensus
Georg S. Seyboth, Dimos V. Dimarogonas, and Karl Henrik Johansson · 2013
Earlier work this paper cites.
Iterative parameter mixing for distributed large-margin training of structured predictors for natural language processing
Gregory F. Coppola · 2015
Earlier work this paper cites.
Dynamic triggering mechanisms for event-triggered control
Antoine Girard · 2015
Earlier work this paper cites.
Distributed convex optimization via continuous-time coordination algorithms with discrete-time communication
Solmaz S. Kia, Jorge Cortés, and Sonia Martínez · 2015
Earlier work this paper cites.
Scalable distributed DNN training using commodity GPU cloud computing
Nikko Strom · 2015
Earlier work this paper cites.
Event-triggered zero-gradient-sum distributed consensus optimization over directed networks
Weisheng Chen and Wei Ren · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Federated learning: Strategies for improving communication efficiency
Jakub Konečnỳ, H Brendan McMahan, Felix X Yu, Peter Richtárik, Ananda Theertha Suresh, and Dave Bacon · 2016
Earlier work this paper cites.
Learning structured sparsity in deep neural networks
Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2016
Earlier work this paper cites.
QSGD: communication-efficient SGD via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Earlier work this paper cites.
Sparse communication for distributed gradient descent
Alham Fikri Aji and Kenneth Heafield · 2017
Earlier work this paper cites.
Asynchronous periodic event-triggered coordination of multi-agent systems
Yaohua Liu, Cameron Nowzari, Zhi Tian, and Qing Ling · 2017
Earlier work this paper cites.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Xiangru Lian, Ce Zhang, Huan Zhang, Cho-Jui Hsieh, Wei Zhang, and Ji Liu · 2017
Cited alongside, same era.
Asynchronous decentralized parallel stochastic gradient descent
Xiangru Lian, Wei Zhang, Ce Zhang, and Ji Liu · 2017
Cited alongside, same era.
Distributed mean estimation with limited communication
A. Theertha Suresh, F. X. Yu, S. Kumar, and H. B. McMahan · 2017
Cited alongside, same era.
Non-convex distributed optimization
Tatiana Tatarenko and Behrouz Touri · 2017
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nathan Srebro, and Benjamin Recht · 2017
Cited alongside, same era.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
A unified analysis of stochastic momentum methods for deep learning
Yan Yan, Tianbao Yang, Zhe Li, Qihang Lin, and Yi Yang · 2018
Later among the works it cites.
Stochastic gradient push for distributed deep learning
Mahmoud Assran, Nicolas Loizou, Nicolas Ballas, and Michael Rabbat · 2019
Later among the works it cites.
Qsparse-local-SGD: Distributed SGD with quantization, sparsification and local computations
Debraj Basu, Deepesh Data, Can Karakus, and Suhas N. Diggavi · 2019
Later among the works it cites.
Error feedback fixes signsgd and other gradient compression schemes
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian U. Stich, and Martin Jaggi · 2019
Later among the works it cites.
Decentralized Stochastic Optimization and Gossip Algorithms with Compressed Communication
Anastasia Koloskova, Sebastian U. Stich, and Martin Jaggi · 2019
Later among the works it cites.
Local SGD Converges Fast and Communicates Little
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li · 2017
Cited alongside, same era.
The convergence of sparsified gradient methods
Dan Alistarh, Torsten Hoefler, Mikael Johansson, Nikola Konstantinov, Sarit Khirirat, and Cédric Renggli · 2018
Cited alongside, same era.
signSGD: Compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Anima Anandkumar · 2018
Cited alongside, same era.
Lag: Lazily aggregated gradient for communication-efficient distributed learning
Tianyi Chen, Georgios Giannakis, Tao Sun, and Wotao Yin · 2018
Cited alongside, same era.
Distributed optimization with dynamic event-triggered mechanisms
Wen Du, Xinlei Yi, Jemin George, Karl Henrik Johansson, and Tao Yang · 2018
Cited alongside, same era.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Y. Lin, S. Han, H. Mao, Y. Wang, and W. J. Dally · 2018
Cited alongside, same era.
Quantized decentralized consensus optimization
Amirhossein Reisizadeh, Aryan Mokhtari, Hamed Hassani, and Ramtin Pedarsani · 2018
Cited alongside, same era.
Sebastian U. Stich · 2019
Later among the works it cites.
Doublesqueeze: Parallel stochastic gradient descent with double-pass error-compensated compression
Hanlin Tang, Chen Yu, Xiangru Lian, Tong Zhang, and Ji Liu · 2019
Later among the works it cites.
Matcha: Speeding up decentralized sgd via matching decomposition sampling
Jianyu Wang, Anit Kumar Sahu, Zhouyi Yang, Gauri Joshi, and Soummya Kar · 2019
Later among the works it cites.
On the linear speedup analysis of communication efficient momentum SGD for distributed non-convex optimization
Hao Yu, Rong Jin, and Sen Yang · 2019
Later among the works it cites.
Parallel restarted SGD with faster convergence and less communication:demystifying why model averaging works for deep learning
Hao Yu, Sen Yang, and Shenghuo Zhu · 2019
Later among the works it cites.
Communication-efficient distributed blockwise momentum sgd with error-feedback
Shuai Zheng, Ziyue Huang, and James Kwok · 2019
Later among the works it cites.
A unified theory of decentralized SGD with changing topology and local updates
Anastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi, and Sebastian U. Stich · 2020
Closest in time.
Decentralized Deep Learning with Arbitrary Communication Compression
Anastasia Koloskova, Tao Lin, Sebastian U. Stich, and Martin Jaggi · 2020
Closest in time.
SPARQ-SGD: Event-triggered and compressed communication in decentralized optimization
Navjot Singh, Deepesh Data, Jemin George, and Suhas Diggavi · 2020
Closest in time.
SlowMo: Improving communication-efficient distributed sgd with slow momentum
Jianyu Wang, Vinayak Tantia, Nicolas Ballas, and Michael Rabbat · 2020
Closest in time.