Fetching the paper…
Reading the bibliography…
Stochastic Gradient Descent (SGD) is a fundamental algorithm in machine learning, representing the optimization backbone for training several classic models, from regression to neural networks.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1986
Earlier work this paper cites.
Parallel and distributed computation: numerical methods
Dimitri P Bertsekas and John N Tsitsiklis · 1989
Earlier work this paper cites.
Distributed computing: fundamentals, simulations, and advanced topics
Hagit Attiya and Jennifer Welch · 2004
Earlier work this paper cites.
Slow learners are fast
John Langford, Alexander J Smola, and Martin Zinkevich · 2009
Earlier work this paper cites.
Distributed delayed stochastic optimization
Alekh Agarwal and John C Duchi · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Project adam: Building an efficient and scalable deep learning training system
Trishul M Chilimbi, Yutaka Suzue, Johnson Apacible, and Karthik Kalyanaraman · 2014
Earlier work this paper cites.
Global convergence of stochastic gradient descent for some non-convex matrix problems
Christopher De Sa, Kunle Olukotun, and Christopher Ré · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and application to data-parallel distributed training of speech dnns
F. Seide, H. Fu, L. G. Jasha, and D. Yu · 2014
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Sébastien Bubeck et al · 2015
Cited alongside, same era.
Asynchronous stochastic convex optimization: the noise is in the noise and sgd don’t care
Sorathan Chaturapruek, John C Duchi, and Christopher Ré · 2015
Cited alongside, same era.
Taming the wild: A unified analysis of hogwild-style algorithms
Christopher M De Sa, Ce Zhang, Kunle Olukotun, and Christopher Ré · 2015
Cited alongside, same era.
Asynchronous stochastic convex optimization
John C Duchi, Sorathan Chaturapruek, and Christopher Ré · 2015
Cited alongside, same era.
Why non-blocking operations should be selfish
Joel Gibson and Vincent Gramoli · 2015
Cited alongside, same era.
Asynchronous parallel stochastic gradient for nonconvex optimization
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Later among the works it cites.
Multi-agent optimization in the presence of byzantine adversaries: fundamental limits
Lili Su and Nitin Vaidya · 2016
Later among the works it cites.
Asynchronous non-bayesian learning in the presence of crash failures
Lili Su and Nitin H Vaidya · 2016
Later among the works it cites.
Fault-tolerant multi-agent optimization: optimal iterative distributed algorithms
Lili Su and Nitin H Vaidya · 2016
Later among the works it cites.
Staleness-aware async-sgd for distributed deep learning
Wei Zhang, Suyog Gupta, Xiangru Lian, and Ji Liu · 2016
Later among the works it cites.
Byzantine-tolerant machine learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu · 2015
Cited alongside, same era.
Asynchronous stochastic coordinate descent: Parallelism and convergence properties
Ji Liu and Stephen J Wright · 2015
Cited alongside, same era.
Temporally bounding tso for fence-free asymmetric synchronization
Adam Morrison and Yehuda Afek · 2015
Cited alongside, same era.
Taming the wild: A unified analysis of hogwild-style algorihms
C. M. De Sa, C. Zhang, K. Olukotun, and C. Re · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Asynchrony begets momentum, with an application to deep learning
Ioannis Mitliagkas, Ce Zhang, Stefan Hadjis, and Christopher Ré · 2016
Cited alongside, same era.
High performance parallel stochastic gradient descent in shared memory
Scott Sallinen, Nadathur Satish, Mikhail Smelyanskiy, Samantika S Sury, and Christopher Ré · 2016
Cited alongside, same era.
Peva Blanchard, El Mahdi El Mhamdi, Rachid Guerraoui, and Julien Stainer · 2017
Later among the works it cites.
Distributed statistical machine learning in adversarial settings: Byzantine gradient descent
Yudong Chen, Lili Su, and Jiaming Xu · 2017
Later among the works it cites.
Byrdie: Byzantine-resilient distributed coordinate descent for decentralized learning
Zhixiong Yang and Waheed U Bajwa · 2017
Later among the works it cites.
Yellowfin and the art of momentum tuning
Jian Zhang, Ioannis Mitliagkas, and Christopher Ré · 2017
Later among the works it cites.
Asynchronous stochastic gradient descent with delay compensation
Shuxin Zheng, Qi Meng, Taifeng Wang, Wei Chen, Nenghai Yu, Zhi-Ming Ma, and Tie-Yan Liu · 2017
Later among the works it cites.
Improved asynchronous parallel optimization analysis for stochastic incremental methods
Rémi Leblond, Fabian Pederegosa, and Simon Lacoste-Julien · 2018
Closest in time.
Sgd and hogwild! convergence without the bounded gradients assumption
Lam M Nguyen, Phuong Ha Nguyen, Marten van Dijk, Peter Richtárik, Katya Scheinberg, and Martin Takáč · 2018
Closest in time.