Fetching the paper…
Reading the bibliography…
Distributed stochastic gradient descent (SGD) algorithms are widely deployed in training large-scale deep learning models, while the communication overhead among workers becomes the new system bottleneck.
Communication-efficient distributed blockwise momentum SGD with error-feedback
Shuai Zheng, Ziyue Huang, and James T Kwok · 1911
Earlier work this paper cites.
Acoustical and environmental robustness in automatic speech recognition
Alejandro Acero · 1990
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 1993
Earlier work this paper cites.
The mnist database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Cifar-10 (canadian institute for advanced research)
Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton · 2010
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Lenet-5, convolutional neural networks
Yann LeCun et al · 2015
Earlier work this paper cites.
Scalable distributed DNN training using commodity GPU cloud computing
Nikko Strom · 2015
Earlier work this paper cites.
Communication quantization for data-parallel training of deep neural networks
Nikoli Dryden, Tim Moon, Sam Ade Jacobs, and Brian Van Essen · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Sparse communication for distributed gradient descent
Alham Fikri Aji and Kenneth Heafield · 2017
Cited alongside, same era.
QSGD: Communication-efficient SGD via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
Inception-v4, inception-resnet and the impact of residual connections on learning
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi · 2017
Cited alongside, same era.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Horovod: fast and easy distributed deep learning in TensorFlow
Alexander Sergeev and Mike Del Balso · 2018
Later among the works it cites.
Efficient top-k query processing on massively parallel hardware
Anil Shanbhag, Holger Pirk, and Samuel Madden · 2018
Later among the works it cites.
Sparsified SGD with memory
Sebastian U Stich, Jean-Baptiste Cordonnier, and Martin Jaggi · 2018
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang · 2018
Later among the works it cites.
Error compensated quantized SGD and its applications to large-scale distributed optimization
Jiaxiang Wu, Weidong Huang, Junzhou Huang, and Tong Zhang · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Cited alongside, same era.
The convergence of sparsified gradient methods
Dan Alistarh, Torsten Hoefler, Mikael Johansson, Nikola Konstantinov, Sarit Khirirat, and Cédric Renggli · 2018
Cited alongside, same era.
AdaComp: Adaptive residual gradient compression for data-parallel distributed training
Chia-Yu Chen, Jungwook Choi, Daniel Brand, Ankur Agrawal, Wei Zhang, and Kailash Gopalakrishnan · 2018
Cited alongside, same era.
Xianyan Jia, Shutao Song, Wei He, Yangzihao Wang, Haidong Rong, Feihu Zhou, Liqiang Xie, Zhenyu Guo, Yuanzhou Yang, Liwei Yu, et al · 2018
Cited alongside, same era.
A linear speedup analysis of distributed deep learning with sparse and quantized communication
Peng Jiang and Gagan Agrawal · 2018
Cited alongside, same era.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and William J Dally · 2018
Cited alongside, same era.
Mixed precision training
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al · 2018
Cited alongside, same era.
Jiarui Fang, Haohuan Fu, Guangwen Yang, and Cho-Jui Hsieh · 2019
Closest in time.
Trading redundancy for communication: Speeding up distributed SGD for non-convex optimization
Farzin Haddadpour, Mohammad Mahdi Kamani, Mehrdad Mahdavi, and Viveck Cadambe · 2019
Closest in time.
Error feedback fixes signsgd and other gradient compression schemes
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian Stich, and Martin Jaggi · 2019
Closest in time.
A distributed synchronous SGD algorithm with global Top- k k sparsification for low bandwidth networks
Shaohuai Shi, Qiang Wang, Kaiyong Zhao, Zhenheng Tang, Yuxin Wang, Xiang Huang, and Xiaowen Chu · 2019
Closest in time.
Peng Sun, Wansen Feng, Ruobing Han, Shengen Yan, and Yonggang Wen · 2019
Closest in time.
Doublesqueeze: Parallel stochastic gradient descent with double-pass error-compensated compression
Hanlin Tang, Chen Yu, Xiangru Lian, Tong Zhang, and Ji Liu · 2019
Closest in time.