Fetching the paper…
Reading the bibliography…
Decentralized training of deep learning models is a key element for enabling data privacy and on-device learning over networks, as well as for efficient scaling to large compute clusters.
Deep South
A. Davis, B. B. Gardner, and M. R. Gardner · 1941
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Gossip-based computation of aggregate information
David Kempe, Alin Dobra, and Johannes Gehrke · 2003
Earlier work this paper cites.
Fast linear iterations for distributed averaging
Lin Xiao and Stephen Boyd · 2004
Earlier work this paper cites.
A scheme for robust distributed sensor fusion based on average consensus
L. Xiao, S. Boyd, and S. Lall · 2005
Earlier work this paper cites.
Randomized gossip algorithms
Stephen Boyd, Arpita Ghosh, Balaji Prabhakar, and Devavrat Shah · 2006
Earlier work this paper cites.
Average consensus on networks with transmission noise or quantization
R. Carli, F. Fagnani, P. Frasca, T. Taylor, and S. Zampieri · 2007
Earlier work this paper cites.
Distributed subgradient methods and quantization effects
Angelia Nedić, Alex Olshevsky, Asuman Ozdaglar, and John N. Tsitsiklis · 2008
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Gossip algorithms for distributed signal processing
A. G. Dimakis, S. Kar, J. M. F. Moura, M. G. Rabbat, and A. Scaglione · 2010
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir, and Lin Xiao · 2012
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2012
Earlier work this paper cites.
Efficiency of coordinate descent methods on huge-scale optimization problems
Yurii Nesterov · 2012
Earlier work this paper cites.
Distributed dual averaging method for multi-agent optimization with quantized communication
Deming Yuan, Shengyuan Xu, Huanyu Zhao, and Lina Rong · 2012
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Scalable distributed DNN training using commodity GPU cloud computing
Nikko Strom · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Federated learning of deep networks using model averaging
H. Brendan McMahan, Eider Moore, Daniel Ramage, and Blaise Agüera y Arcas · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Cited alongside, same era.
Distributed coordinate descent method for learning with big data
Peter Richtárik and Martin Takáč · 2016
Cited alongside, same era.
Parallel SGD: When does averaging help?
Jian Zhang, Christopher De Sa, Ioannis Mitliagkas, and Christopher Ré · 2016
Cited alongside, same era.
QSGD: Communication-efficient SGD via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
Accurate, large minibatch SGD: Training ImageNet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Sparsified SGD with memory
Sebastian U Stich, Jean-Baptiste Cordonnier, and Martin Jaggi · 2018
Later among the works it cites.
Communication compression for decentralized training
Hanlin Tang, Shaoduo Gan, Ce Zhang, Tong Zhang, and Ji Liu · 2018
Later among the works it cites.
Jianyu Wang and Gauri Joshi · 2018
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang · 2018
Later among the works it cites.
Stochastic gradient push for distributed deep learning
Mahmoud Assran, Nicolas Loizou, Nicolas Ballas, and Mike Rabbat · 2019
Closest in time.
Nested distributed gradient methods with adaptive quantized communication
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Xiangru Lian, Ce Zhang, Huan Zhang, Cho-Jui Hsieh, Wei Zhang, and Ji Liu · 2017
Cited alongside, same era.
Communication-Efficient Learning of Deep Networks from Decentralized Data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Cited alongside, same era.
Regularizing and optimizing LSTM language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2017
Cited alongside, same era.
Optimal algorithms for smooth and strongly convex distributed optimization in networks
Kevin Scaman, Francis Bach, Sébastien Bubeck, Yin Tat Lee, and Laurent Massoulié · 2017
Cited alongside, same era.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Cited alongside, same era.
The convergence of sparsified gradient methods
Dan Alistarh, Torsten Hoefler, Mikael Johansson, Nikola Konstantinov, Sarit Khirirat, and Cedric Renggli · 2018
Cited alongside, same era.
signSGD: Compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar · 2018
Cited alongside, same era.
Albert S. Berahas, Charikleia Iakovidou, and Ermin Wei · 2019
Closest in time.
Stochastic distributed learning with gradient quantization and variance reduction
Samuel Horváth, Dmitry Kovalev, Konstantin Mishchenko, Peter Richtárik, and Sebastian Urban Stich · 2019
Closest in time.
Advances and open problems in federated learning
Peter Kairouz, H. Brendan McMahan, and et. al. including Sebastian U. Stich · 2019
Closest in time.
Error feedback fixes SignSGD and other gradient compression schemes
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian Stich, and Martin Jaggi · 2019
Closest in time.
Decentralized stochastic optimization and gossip algorithms with compressed communication
Anastasia Koloskova, Sebastian Stich, and Martin Jaggi · 2019
Closest in time.
Distributed learning with compressed gradient differences
Konstantin Mishchenko, Eduard Gorbunov, Martin Takáč, and Peter Richtárik · 2019
Closest in time.
Robust and communication-efficient collaborative learning
Amirhossein Reisizadeh, Hossein Taheri, Aryan Mokhtari, Hamed Hassani, and Ramtin Pedarsani · 2019
Closest in time.
Local sgd converges fast and communicates little
Sebastian U. Stich · 2019
Closest in time.
Sebastian U Stich and Sai Praneeth Karimireddy · 2019
Closest in time.
Deepsqueeze: Decentralization meets error-compensated compression
Hanlin Tang, Xiangru Lian, Shuang Qiu, Lei Yuan, Ce Zhang, Tong Zhang, and Ji Liu · 2019
Closest in time.
MATCHA: speeding up decentralized SGD via matching decomposition sampling
Jianyu Wang, Anit Kumar Sahu, Zhouyi Yang, Gauri Joshi, and Soummya Kar · 2019
Closest in time.
On the linear speedup analysis of communication efficient momentum SGD for distributed non-convex optimization
Hao Yu, Rong Jin, and Sen Yang · 2019
Closest in time.
Don’t use large mini-batches, use local SGD
Tao Lin, Sebastian U. Stich, Kumar Kshitij Patel, and Martin Jaggi · 2020
Closest in time.