Fetching the paper…
Reading the bibliography…
Training neural networks on large datasets can be accelerated by distributing the workload over a network of machines.
Sui confini della probabilità
Francesco Paolo Cantelli · 1928
Earlier work this paper cites.
A Stochastic Approximation Method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
J J Hopfield · 1982
Earlier work this paper cites.
A Direct Adaptive Method for Faster Backpropagation Learning: the RPROP Algorithm
M. Riedmiller and H. Braun · 1993
Earlier work this paper cites.
The Three Sigma Rule
Friedrich Pukelsheim · 1994
Earlier work this paper cites.
Cubic Regularization of Newton Method and its Global Performance
Yurii Nesterov and B.T. Polyak · 2006
Earlier work this paper cites.
1-Bit Stochastic Gradient Descent and Application to Data-Parallel Distributed Training of Speech DNNs
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Stochastic spectral descent for discrete graphical models
David Carlson, Ya-Ping Hsieh, Edo Collins, Lawrence Carin, and Volkan Cevher · 2016
Cited alongside, same era.
QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
Natasha 2: Faster Non-Convex Optimization Than SGD
Zeyuan Allen-Zhu · 2017
Cited alongside, same era.
Machine Learning with Adversaries: Byzantine Tolerant Gradient Descent
Peva Blanchard, El Mahdi El Mhamdi, Rachid Guerraoui, and Julien Stainer · 2017
Cited alongside, same era.
Quasi-Recurrent Neural Networks
James Bradbury, Stephen Merity, Caiming Xiong, and Richard Socher · 2017
Cited alongside, same era.
Automatic Differentiation in PyTorch
signSGD: Compressed Optimisation for Non-Convex Problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar · 2018
Closest in time.
The Hidden Vulnerability of Distributed Learning in Byzantium
El Mahdi El Mhamdi, Rachid Guerraoui, and Sébastien Rouault · 2018
Closest in time.
Gloo Collective Communications Library, 2018
Gloo · 2018
Closest in time.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and Bill Dally · 2018
Closest in time.
Nvidia Collective Communications Library, 2018
NCCL · 2018
Closest in time.
On the Convergence of Adam and Beyond
Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
TernGrad: Ternary Gradients to Reduce Communication in Distributed Deep Learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Cited alongside, same era.
Byzantine Stochastic Gradient Descent
Dan Alistarh, Zeyuan Allen-Zhu, and Jerry Li · 2018
Cited alongside, same era.
Dissecting Adam: The Sign, Magnitude and Variance of Stochastic Gradients
Lukas Balles and Philipp Hennig · 2018
Cited alongside, same era.
Theoria combinationis observationum erroribus minimis obnoxiae, pars prior
Carl Friedrich Gauss
Cited in the paper.
IBM Summit Supercomputer, 2018
TOP500 · 2018
Closest in time.
ATOMO: Communication-efficient Learning via Atomic Sparsification
Hongyi Wang, Scott Sievert, Shengchao Liu, Zachary B. Charles, Dimitris S. Papailiopoulos, and Stephen Wright · 2018
Closest in time.
Byzantine-Robust Distributed Learning: Towards Optimal Statistical Rates
Dong Yin, Yudong Chen, Ramchandran Kannan, and Peter Bartlett · 2018
Closest in time.