Fetching the paper…
Reading the bibliography…
Running Stochastic Gradient Descent (SGD) in a decentralized fashion has shown promising results.
Using MPI-2: Advanced features of the message passing interface
William Gropp, Rajeev Thakur, and Ewing Lusk · 1999
Earlier work this paper cites.
Certificateless public key cryptography
Sattam S Al-Riyami and Kenneth G Paterson · 2003
Earlier work this paper cites.
Solving large scale linear prediction problems using stochastic gradient descent algorithms
Tong Zhang · 2004
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al · 2012
Earlier work this paper cites.
Random search for hyper-parameter optimization
James Bergstra and Yoshua Bengio · 2012
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
The cifar-10 dataset
Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems
Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang · 2015
Earlier work this paper cites.
Decentralized double stochastic averaging gradient
Aryan Mokhtari and Alejandro Ribeiro · 2015
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally · 2015
Earlier work this paper cites.
Deep learning with limited numerical precision
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan · 2015
Earlier work this paper cites.
Tensorflow: a system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Earlier work this paper cites.
Cntk: Microsoft’s open-source deep-learning toolkit
Frank Seide and Amit Agarwal · 2016
Earlier work this paper cites.
Consensus optimization with delayed and stochastic gradients on decentralized networks
Benjamin Sirb and Xiaojing Ye · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Zipml: Training linear models with end-to-end low precision, and a little bit of deep learning
Hantian Zhang, Jerry Li, Kaan Kara, Dan Alistarh, Ji Liu, and Ce Zhang · 2017
Cited alongside, same era.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
Gradient sparsification for communication-efficient distributed optimization
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang · 2018
Later among the works it cites.
High-accuracy low-precision training
Christopher De Sa, Megan Leszczynski, Jian Zhang, Alana Marzoev, Christopher R Aberger, Kunle Olukotun, and Christopher Ré · 2018
Later among the works it cites.
Cola: Decentralized linear learning
Lie He, An Bian, and Martin Jaggi · 2018
Later among the works it cites.
Stochastic gradient push for distributed deep learning
Mahmoud Assran, Nicolas Loizou, Nicolas Ballas, and Michael Rabbat · 2018
Later among the works it cites.
Sparsified sgd with memory
Sebastian U Stich, Jean-Baptiste Cordonnier, and Martin Jaggi · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Cited alongside, same era.
Communication-efficient algorithms for decentralized and stochastic optimization
Guanghui Lan, Soomin Lee, and Yi Zhou · 2017
Cited alongside, same era.
Distributed mean estimation with limited communication
Ananda Theertha Suresh, Felix X Yu, Sanjiv Kumar, and H Brendan McMahan · 2017
Cited alongside, same era.
Understanding and optimizing asynchronous low-precision stochastic gradient descent
Christopher De Sa, Matthew Feldman, Christopher Ré, and Kunle Olukotun · 2017
Cited alongside, same era.
Training quantized nets: A deeper understanding
Hao Li, Soham De, Zheng Xu, Christoph Studer, Hanan Samet, and Tom Goldstein · 2017
Cited alongside, same era.
Markov chains and mixing times , volume 107
David A Levin and Yuval Peres · 2017
Cited alongside, same era.
A brief tutorial on distributed and concurrent machine learning
Dan Alistarh · 2018
Cited alongside, same era.
Dan Alistarh, Torsten Hoefler, Mikael Johansson, Nikola Konstantinov, Sarit Khirirat, and Cédric Renggli · 2018
Later among the works it cites.
Synchronous multi-gpu deep learning with low-precision communication: An experimental study
D Grubic, L Tam, Dan Alistarh, and Ce Zhang · 2018
Later among the works it cites.
A linear speedup analysis of distributed deep learning with sparse and quantized communication
Peng Jiang and Gagan Agrawal · 2018
Later among the works it cites.
Quantized decentralized consensus optimization
Amirhossein Reisizadeh, Aryan Mokhtari, S. Hamed Hassani, and Ramtin Pedarsani · 2018
Later among the works it cites.
Local sgd converges fast and communicates little
Sebastian U Stich · 2018
Later among the works it cites.
Deepsqueeze: Parallel stochastic gradient descent with double-pass error-compensated compression
Hanlin Tang, Xiangru Lian, Shuang Qiu, Lei Yuan, Ce Zhang, Tong Zhang, and Ji Liu · 2019
Later among the works it cites.
Decentralized stochastic optimization and gossip algorithms with compressed communication
Anastasia Koloskova, Sebastian U Stich, and Martin Jaggi · 2019
Later among the works it cites.
Dadam: A consensus-based distributed adaptive gradient method for online optimization
Parvin Nazari, Davoud Ataee Tarzanagh, and George Michailidis · 2019
Later among the works it cites.
Asynchronous decentralized optimization in directed networks
Jiaqi Zhang and Keyou You · 2019
Later among the works it cites.
Distributed learning with sublinear communication
Jayadev Acharya, Christopher De Sa, Dylan J Foster, and Karthik Sridharan · 2019
Later among the works it cites.