Fetching the paper…
Reading the bibliography…
Modern deep learning models are often trained in parallel over a collection of distributed machines to reduce training time.
Television by pulse code modulation
W. M. Goodall · 1951
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Picture coding using pseudo-random noise
L. Roberts · 1962
Earlier work this paper cites.
Stabilization of linear systems with limited information
Nicola Elia and Sanjoy K Mitter · 2001
Earlier work this paper cites.
Efficient large-scale distributed training of conditional maximum entropy models
Ryan McDonald, Mehryar Mohri, Nathan Silberman, Dan Walker, and Gideon S. Mann · 2009
Earlier work this paper cites.
Scaling up machine learning: Parallel and distributed approaches
Ron Bekkerman, Mikhail Bilenko, and John Langford · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Scalar quantization for relative error
John Z Sun and Vivek K Goyal · 2011
Earlier work this paper cites.
A framework for bayesian optimality of psychophysical laws
John Z Sun, Grace I Wang, Vivek K Goyal, and Lav R Varshney · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course
Yurii Nesterov · 2013
Earlier work this paper cites.
Fast distributed coordinate descent for minimizing non-strongly convex losses
Olivier Fercoq, Zheng Qu, Peter Richtárik, and Martin Takáč · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Communication-efficient distributed optimization using an approximate Newton-type method
Ohad Shamir, Nati Srebro, and Tong Zhang · 2014
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Sébastien Bubeck et al · 2015
Earlier work this paper cites.
Deep learning with limited numerical precision
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan · 2015
Earlier work this paper cites.
ADAM: A Method for Stochastic Optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Communication quantization for data-parallel training of deep neural networks
N. Dryden, T. Moon, S. A. Jacobs, and B. V. Essen · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Sparse communication for distributed gradient descent
Alham Fikri Aji and Kenneth Heafield · 2017
Cited alongside, same era.
QSGD: Communication-efficient sgd via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
Accurate, large minibatch SGD: training ImageNet in 1 hour
Priya Goyal, Piotr Dollár, Ross B. Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Distributed learning with compressed gradients
Sarit Khirirat, Hamid Reza Feyzmahdavian, and Mikael Johansson · 2018
Later among the works it cites.
Randomized distributed mean estimation: accuracy vs communication
Jakub Konečný and Peter Richtárik · 2018
Later among the works it cites.
3lc: Lightweight and effective traffic compression for distributed machine learning
Hyeontaek Lim, David G Andersen, and Michael Kaminsky · 2018
Later among the works it cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and Bill Dally · 2018
Later among the works it cites.
Local SGD converges fast and communicates little
Sebastian U. Stich · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Neural collaborative filtering
Xiangnan He, Lizi Liao, Hanwang Zhang, Liqiang Nie, Xia Hu, and Tat-Seng Chua · 2017
Cited alongside, same era.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger · 2017
Cited alongside, same era.
Quantized neural networks: Training neural networks with low precision weights and activations
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2017
Cited alongside, same era.
On-chip training of recurrent neural networks with limited numerical precision
T. Na, J. H. Ko, J. Kung, and S. Mukhopadhyay · 2017
Cited alongside, same era.
Distributed mean estimation with limited communication
Ananda Theertha Suresh, Felix X. Yu, Sanjiv Kumar, and H. Brendan McMahan · 2017
Cited alongside, same era.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Cited alongside, same era.
Sparsified SGD with memory
Sebastian U Stich, Jean-Baptiste Cordonnier, and Martin Jaggi · 2018
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang · 2018
Later among the works it cites.
Nonconvex variance reduced optimization with arbitrary sampling
Samuel Horváth and Peter Richtárik · 2019
Closest in time.
Stochastic distributed learning with gradient quantization and variance reduction
Samuel Horváth, Dmitry Kovalev, Konstantin Mishchenko, Sebastian Stich, and Peter Richtárik · 2019
Closest in time.
Error feedback fixes SignSGD and other gradient compression schemes
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian U. Stich, and Martin Jaggi · 2019
Closest in time.
Distributed learning with compressed gradient differences
Konstantin Mishchenko, Eduard Gorbunov, Martin Takáč, and Peter Richtárik · 2019
Closest in time.
DoubleSqueeze
Hanlin Tang, Chen Yu, Xiangru Lian, Tong Zhang, and Ji Liu · 2019
Closest in time.
Powersgd: Practical low-rank gradient compression for distributed optimization
Thijs Vogels, Sai Praneeth Karimireddy, and Martin Jaggi · 2019
Closest in time.
Communication-efficient distributed blockwise momentum sgd with error-feedback
Shuai Zheng, Ziyue Huang, and James Kwok · 2019
Closest in time.
Efficient Sparse Collective Communication and its application to Accelerate Distributed Deep Learning
Jiawei Fei, Chen-Yu Ho, Atal Narayan Sahu, Marco Canini, and Amedeo Sapio · 2020
Closest in time.
Scaling Distributed Machine Learning with In-Network Aggregation
Amedeo Sapio, Marco Canini, Chen-Yu Ho, Jacob Nelson, Panos Kalnis, Changhoon Kim, Arvind Krishnamurthy, Masoud Moshref, Dan R. K. Ports, and Peter Richtárik · 2021
Closest in time.