Fetching the paper…
Reading the bibliography…
This paper studies the problem of error-runtime trade-off, typically encountered in decentralized training based on stochastic gradient descent (SGD) using a given network.
Distributed asynchronous deterministic and stochastic gradient optimization algorithms
John Tsitsiklis, Dimitri Bertsekas, and Michael Athans · 1986
Earlier work this paper cites.
A constructive proof of vizing’s theorem
Jayadev Misra and David Gries · 1992
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz · 1993
Earlier work this paper cites.
Fast linear iterations for distributed averaging
Lin Xiao and Stephen Boyd · 2004
Earlier work this paper cites.
Sensor networks with random links: Topology design for distributed consensus
Soummya Kar and José MF Moura · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Distributed subgradient methods for multi-agent optimization
Angelia Nedic and Asuman Ozdaglar · 2009
Earlier work this paper cites.
Dual averaging for distributed optimization: Convergence analysis and network scaling
John C Duchi, Alekh Agarwal, and Martin J Wainwright · 2012
Earlier work this paper cites.
Communication/computation tradeoffs in consensus-based distributed optimization
Konstantinos Tsianos, Sean Lawlor, and Michael G Rabbat · 2012
Earlier work this paper cites.
Modern graph theory
Béla Bollobás · 2013
Earlier work this paper cites.
Expander graph and communication-efficient decentralized optimization
Yat-Tin Chow, Wei Shi, Tianyu Wu, and Wotao Yin · 2016
Earlier work this paper cites.
Model accuracy and runtime tradeoff in distributed deep learning: A systematic study
Suyog Gupta, Wei Zhang, and Fei Wang · 2016
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
H Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, et al · 2016
Cited alongside, same era.
Using the output embedding to improve language models
Ofir Press and Lior Wolf · 2016
Cited alongside, same era.
Excess-risk of distributed stochastic learners
Zaid J Towfic, Jianshu Chen, and Ali H Sayed · 2016
Cited alongside, same era.
On the convergence of decentralized gradient descent
Kun Yuan, Qing Ling, and Wotao Yin · 2016
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Later among the works it cites.
Slow and stale gradients can win the race: Error-runtime trade-offs in distributed SGD
Sanghamitra Dutta, Gauri Joshi, Soumyadip Ghosh, Parijat Dube, and Priya Nagpurkar · 2018
Later among the works it cites.
Convergence rates for distributed stochastic optimization over random networks
Dusan Jakovetic, Dragana Bajovic, Anit Kumar Sahu, and Soummya Kar · 2018
Later among the works it cites.
Network topology and communication-computation tradeoffs in decentralized optimization
Angelia Nedić, Alex Olshevsky, and Michael G Rabbat · 2018
Later among the works it cites.
Optimal algorithms for non-smooth distributed optimization in networks
Kevin Scaman, Francis Bach, Sébastien Bubeck, Laurent Massoulié, and Yin Tat Lee · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jinshan Zeng and Wotao Yin · 2016
Cited alongside, same era.
Collaborative deep learning in fixed topology networks
Zhanhong Jiang, Aditya Balu, Chinmay Hegde, and Soumik Sarkar · 2017
Cited alongside, same era.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Xiangru Lian, Ce Zhang, Huan Zhang, Cho-Jui Hsieh, Wei Zhang, and Ji Liu · 2017
Cited alongside, same era.
Asynchronous decentralized parallel stochastic gradient descent
Xiangru Lian, Wei Zhang, Ce Zhang, and Ji Liu · 2017
Cited alongside, same era.
Stochastic gradient push for distributed deep learning
Mahmoud Assran, Nicolas Loizou, Nicolas Ballas, and Michael Rabbat · 2018
Cited alongside, same era.
Sebastian U Stich · 2018
Later among the works it cites.
Communication compression for decentralized training
Hanlin Tang, Shaoduo Gan, Ce Zhang, Tong Zhang, and Ji Liu · 2018
Later among the works it cites.
Adaptive communication strategies to achieve the best error-runtime trade-off in local-update SGD
Jianyu Wang and Gauri Joshi · 2018
Later among the works it cites.
Jianyu Wang and Gauri Joshi · 2018
Later among the works it cites.
Decentralized stochastic optimization and gossip algorithms with compressed communication
Anastasia Koloskova, Sebastian U Stich, and Martin Jaggi · 2019
Closest in time.