Fetching the paper…
Reading the bibliography…
Communication overhead is one of the key challenges that hinders the scalability of distributed optimization algorithms.
Distributed training strategies for the structured perceptron
Ryan McDonald, Keith Hall, and Gideon Mann · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J Smola · 2010
Earlier work this paper cites.
Distributed delayed stochastic optimization
Alekh Agarwal and John C Duchi · 2011
Earlier work this paper cites.
Hybrid deterministic-stochastic methods for data fitting
Michael P Friedlander and Mark Schmidt · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Communication-efficient algorithms for statistical optimization
Yuchen Zhang, Martin J Wainwright, and John C Duchi · 2012
Earlier work this paper cites.
Parallel training of dnns with natural gradient and parameter averaging
Daniel Povey, Xiaohui Zhang, and Sanjeev Khudanpur · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Distributed stochastic optimization and learning
Ohad Shamir and Nathan Srebro · 2014
Earlier work this paper cites.
Taming the wild: A unified analysis of hogwild-style algorithms
Christopher M De Sa, Ce Zhang, Kunle Olukotun, and Christopher Ré · 2015
Earlier work this paper cites.
Deep learning with limited numerical precision
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan · 2015
Earlier work this paper cites.
Asynchronous parallel stochastic gradient for nonconvex optimization
Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu · 2015
Earlier work this paper cites.
Scalable distributed dnn training using commodity gpu cloud computing
Nikko Strom · 2015
Earlier work this paper cites.
Experiments on parallel training of deep neural network using model averaging
Hang Su and Haoyu Chen · 2015
Earlier work this paper cites.
Scalable training of deep learning machines by incremental block training with intra-block parallel optimization and blockwise model-update filtering
Kai Chen and Qiang Huo · 2016
Earlier work this paper cites.
Communication quantization for data-parallel training of deep neural networks
Nikoli Dryden, Tim Moon, Sam Ade Jacobs, and Brian Van Essen · 2016
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Earlier work this paper cites.
Cntk: Microsoft’s open-source deep-learning toolkit
Frank Seide and Amit Agarwal · 2016
Cited alongside, same era.
Parallel sgd: When does averaging help?
Jian Zhang, Christopher De Sa, Ioannis Mitliagkas, and Christopher Ré · 2016
Cited alongside, same era.
Sparse communication for distributed gradient descent
Alham Fikri Aji and Kenneth Heafield · 2017
Cited alongside, same era.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
On the rates of convergence of parallelized averaged stochastic gradient algorithms
Antoine Godichon-Baggioni and Sofiane Saadane · 2017
Cited alongside, same era.
Cross-iteration coded computing
Farzin Haddadpour, Yaoqing Yang, Viveck Cadambe, and Pulkit Grover · 2018
Later among the works it cites.
Parallelizing stochastic gradient descent for least squares regression: mini-batching, averaging, and model misspecification
Prateek Jain, Sham M Kakade, Rahul Kidambi, Praneeth Netrapalli, and Aaron Sidford · 2018
Later among the works it cites.
Efficient decentralized deep learning by dynamic model averaging
Michael Kamp, Linara Adilova, Joachim Sicking, Fabian Hüger, Peter Schlicht, Tim Wirtz, and Stefan Wrobel · 2018
Later among the works it cites.
Don’t use large mini-batches, use local sgd
Tao Lin, Sebastian U Stich, and Martin Jaggi · 2018
Later among the works it cites.
Sparsified sgd with memory
Sebastian U Stich, Jean-Baptiste Cordonnier, and Martin Jaggi · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and William J Dally · 2017
Cited alongside, same era.
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2017
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Cited alongside, same era.
On-chip training of recurrent neural networks with limited numerical precision
Taesik Na, Jong Hwan Ko, Jaeha Kung, and Saibal Mukhopadhyay · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
meprop: Sparsified back propagation for accelerated deep learning with reduced overfitting
Xu Sun, Xuancheng Ren, Shuming Ma, and Houfeng Wang · 2017
Cited alongside, same era.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Cited alongside, same era.
Jianyu Wang and Gauri Joshi · 2018
Later among the works it cites.
Jianyu Wang and Gauri Joshi · 2018
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang · 2018
Later among the works it cites.
Improving the privacy and accuracy of ADMM-based distributed algorithms
Xueru Zhang, Mohammad Mahdi Khalili, and Mingyan Liu · 2018
Later among the works it cites.
On the convergence properties of a k-step averaging stochastic gradient descent algorithm for nonconvex optimization
Fan Zhou and Guojing Cong · 2018
Later among the works it cites.
Distributed asynchronous optimization with unbounded delays: How slow can you go?
Zhengyuan Zhou, Panayotis Mertikopoulos, Nicholas Bambos, Peter W Glynn, Yinyu Ye, Li-Jia Li, and Fei-Fei Li · 2018
Later among the works it cites.
Trading redundancy for communication: Speeding up distributed sgd for non-convex optimization
Farzin Haddadpour, Mohammad Mahdi Kamani, Mehrdad Mahdavi, and Viveck Cadambe · 2019
Closest in time.
Local sgd converges fast and communicates little
Sebastian Urban Stich · 2019
Closest in time.
On the computation and communication complexity of parallel sgd with dynamic batch sizes for stochastic non-convex optimization
Hao Yu and Rong Jin · 2019
Closest in time.
Parallel restarted sgd with faster convergence and less communication: Demystifying why model averaging works for deep learning
Hao Yu, Sen Yang, and Shenghuo Zhu · 2019
Closest in time.
Recycled admm: Improving the privacy and accuracy of distributed algorithms
X. Zhang, M. M. Khalili, and M. Liu · 2019
Closest in time.