Fetching the paper…
Reading the bibliography…
Distributed optimization is essential for training large models on large datasets.
Decentralized deep learning with arbitrary communication compression
Anastasia Koloskova, Tao Lin, Sebastian U Stich, and Martin Jaggi · 1907
Earlier work this paper cites.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Z. Li, Ryota Tomioka, and Milan Vojnovic · 2007
Earlier work this paper cites.
TernGrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton · 2009
Earlier work this paper cites.
Distributed training strategies for the structured perceptron
Ryan McDonald, Keith Hall, and Gideon Mann · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Deep learning with elastic averaged SGD
S. Zhang, A. Choromanska, and Y. LeCun · 2015
Earlier work this paper cites.
Revisiting distributed synchronous SGD
Jianmin Chen, Rajat Monga, Samy Bengio, and Rafal Jozefowicz · 2016
Earlier work this paper cites.
Scalable training of deep learning machines by incremental block training with intra-block parallel optimization and blockwise model-update filtering
Kai Chen and Qiang Huo · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Stochastic gradient-push for strongly convex functions on time-varying directed graphs
Angelia Nedić and Alex Olshevsky · 2016
Earlier work this paper cites.
Accurate, large minibatch SGD: Training ImageNet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Collaborative deep learning in fixed topology networks
Zhanhong Jiang, Aditya Balu, Chinmay Hegde, and Soumik Sarkar · 2017
Cited alongside, same era.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Xiangru Lian, Ce Zhang, Huan Zhang, Cho-Jui Hsieh, Wei Zhang, and Ji Liu · 2017
Cited alongside, same era.
Nicolas Loizou and Peter Richtárik · 2017
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
H. Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Agüera y Arcas · 2017
Cited alongside, same era.
Pytorch: Tensors and dynamic neural networks in python with strong gpu acceleration, 2017
Adam Paszke, Soumith Chintala, Ronan Collobert, Koray Kavukcuoglu, Clement Farabet, Samy Bengio, Iain Melvin, Jason Weston, and Johnny Mariethoz · 2017
signSGD with majority vote is communication efficient and fault tolerant
Jeremy Bernstein, Jiawei Zhao, Kamyar Azizzadenesheli, and Anima Anandkumar · 2019
Closest in time.
Accelerated linear convergence of stochastic momentum methods in Wasserstein distances
Bugra Can, Mert Gürbüzbalaban, and Lingjiong Zhu · 2019
Closest in time.
Anytime minibatch: Exploiting stragglers in online distributed optimization
Nuwan Ferdinand, Haider Al-Lawati, Stark Draper, and Matthew Nokelby · 2019
Closest in time.
Understanding the role of momentum in stochastic gradient methods
Igor Gitman, Hunter Lang, Pengchuan Zhang, and Lin Xiao · 2019
Closest in time.
Error feedback fixes SignSGD and other gradient compression schemes
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian Stich, and Martin Jaggi · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E. Curtis, and Jorge Nocedal · 2018
Cited alongside, same era.
Slow and stale gradients can win the race: Error-runtime trade-offs in distributed SGD
Sanghamitra Dutta, Gauri Joshi, Soumyadip Ghosh, Parijat Dube, and Priya Nagpurkar · 2018
Cited alongside, same era.
Asynchronous decentralized parallel stochastic gradient descent
Xiangru Lian, Wei Zhang, Ce Zhang, and Ji Liu · 2018
Cited alongside, same era.
Don’t use large mini-batches, use local sgd
Tao Lin, Sebastian U Stich, Kumar Kshitij Patel, and Martin Jaggi · 2018
Cited alongside, same era.
Scaling neural machine translation
Myle Ott, Grangier David Edunov, Sergey, and Michael Auli · 2018
Cited alongside, same era.
Jianyu Wang and Gauri Joshi · 2018
Cited alongside, same era.
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Closest in time.
Quasi-hyperbolic momentum and Adam for deep learning
Jerry Ma and Denis Yarats · 2019
Closest in time.
Language models are unsupervised multi-task learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Closest in time.
Optimal statistical rates for decentralised non-parametric regression with linear speed-up
Dominic Richards and Patrick Rebeschini · 2019
Closest in time.
Optimal convergence rates for convex distributed optimization in networks
Kevin Scaman, Francis Bach, Sébastien Bubeck, Yin Tat Lee, and Laurent Massoulié · 2019
Closest in time.
Local SGD converges fast and communicates little
Sebastian U Stich · 2019
Closest in time.
PowerSGD: Practical low-rank gradient compression for distributed optimization
Thijs Vogels, Sai Praneeth Karimireddy, and Martin Jaggi · 2019
Closest in time.
MATCHA: Speeding up decentralized SGD via matching decomposition sampling
Jianyu Wang, Anit Kumar Sahu, Zhouyi Yang, Gauri Joshi, and Soummya Kar · 2019
Closest in time.
Lookahead optimizer: k steps forward, 1 step back
Michael R Zhang, James Lucas, Geoffrey Hinton, and Jimmy Ba · 2019
Closest in time.