Fetching the paper…
Reading the bibliography…
Distributed stochastic gradient descent is an important subroutine in distributed learning.
Universal codeword sets and representations of the integers
Peter Elias · 1975
Earlier work this paper cites.
Least squares quantization in PCM
Stuart Lloyd · 1982
Earlier work this paper cites.
An elementary proof of a theorem of johnson and lindenstrauss
Sanjoy Dasgupta and Anupam Gupta · 2003
Earlier work this paper cites.
Approximate nearest neighbors and the fast Johnson-Lindenstrauss transform
Nir Ailon and Bernard Chazelle · 2006
Earlier work this paper cites.
Our data, ourselves: Privacy via distributed noise generation
Cynthia Dwork, Krishnaram Kenthapadi, Frank McSherry, Ilya Mironov, and Moni Naor · 2006
Earlier work this paper cites.
Calibrating noise to sensitivity in private data analysis
Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam Smith · 2006
Earlier work this paper cites.
The infinite mnist dataset, 2007
Leon Bottou · 2007
Earlier work this paper cites.
Distributed training strategies for the structured perceptron
Ryan McDonald, Keith Hall, and Gideon Mann · 2010
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al · 2012
Earlier work this paper cites.
Hadamard matrices and their applications
Kathy J Horadam · 2012
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, Karthik Sridharan, et al · 2012
Earlier work this paper cites.
Deep learning with cots hpc systems
Adam Coates, Brody Huval, Tao Wang, David Wu, Bryan Catanzaro, and Ng Andrew · 2013
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Cited alongside, same era.
Signal processing and machine learning with differential privacy: Algorithms and challenges for continuous data
Anand D Sarwate and Kamalika Chaudhuri · 2013
Cited alongside, same era.
Private empirical risk minimization: Efficient algorithms and tight error bounds
Raef Bassily, Adam Smith, and Abhradeep Thakurta · 2014
Cited alongside, same era.
The algorithmic foundations of differential privacy
Cynthia Dwork and Aaron Roth · 2014
Cited alongside, same era.
Scaling distributed machine learning with the parameter server
Mu Li, David G Andersen, Jun Woo Park, Alexander J Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J Shekita, and Bor-Yiing Su · 2014
Cited alongside, same era.
Communication efficient distributed machine learning with the parameter server
Federated learning: Strategies for improving communication efficiency
Jakub Konečnỳ, H Brendan McMahan, Felix X Yu, Peter Richtárik, Ananda Theertha Suresh, and Dave Bacon · 2016
Later among the works it cites.
Randomized distributed mean estimation: Accuracy vs communication
Jakub Konečnỳ and Peter Richtárik · 2016
Later among the works it cites.
Communication-efficient learning of deep networks from decentralized data
H. Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2016
Later among the works it cites.
Federated learning of deep networks using model averaging
H. Brendan McMahan, Eider Moore, Daniel Ramage, and Blaise Aguera y Arcas · 2016
Later among the works it cites.
Communication-efficient stochastic gradient descent, with applications to neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mu Li, David G Andersen, Alexander J Smola, and Kai Yu · 2014
Cited alongside, same era.
Parallel training of deep neural networks with natural gradient and parameter averaging
Daniel Povey, Xiaohui Zhang, and Sanjeev Khudanpur · 2014
Cited alongside, same era.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Cited alongside, same era.
Recurrent neural network regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals · 2014
Cited alongside, same era.
Deep learning with limited numerical precision
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan · 2015
Cited alongside, same era.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al · 2016
Cited alongside, same era.
Deep learning with differential privacy
Martín Abadi, Andy Chu, Ian Goodfellow, H Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang · 2016
Cited alongside, same era.
Dan Alistarh, Demjan Grubic, Jerry Liu, Ryota Tomioka, and Milan Vojnovic · 2017
Later among the works it cites.
Practical secure aggregation for privacy-preserving machine learning
Keith Bonawitz, Vladimir Ivanov, Ben Kreuter, Antonio Marcedone, H Brendan McMahan, Sarvar Patel, Daniel Ramage, Aaron Segal, and Karn Seth · 2017
Later among the works it cites.
The composition theorem for differential privacy
Peter Kairouz, Sewoong Oh, and Pramod Viswanath · 2017
Later among the works it cites.
Distributed mean estimation with limited communication
Ananda Theertha Suresh, X Yu Felix, Sanjiv Kumar, and H Brendan McMahan · 2017
Later among the works it cites.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Later among the works it cites.
Bolt-on differential privacy for scalable stochastic gradient descent-based analytics
Xi Wu, Fengan Li, Arun Kumar, Kamalika Chaudhuri, Somesh Jha, and Jeffrey Naughton · 2017
Later among the works it cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and Bill Dally · 2018
Closest in time.
Variance-based gradient compression for efficient distributed deep learning, 2018
Takuya Akiba Yusuke Tsuzuku, Hiroto Imachi · 2018
Closest in time.