A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Distributed asynchronous deterministic and stochastic gradient optimization algorithms
John Tsitsiklis, Dimitri Bertsekas, and Michael Athans · 1986
Earlier work this paper cites.
A course in microeconomic theory
David M Kreps · 1990
Earlier work this paper cites.
The mnist database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
A first course in probability
Ross Sheldon · 2002
Earlier work this paper cites.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Distributed delayed stochastic optimization
Alekh Agarwal and John C Duchi · 2011
Earlier work this paper cites.
Algorithms
Robert Sedgewick and Kevin Wayne · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean et al · 2012
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir, and Lin Xiao · 2012
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence rate for finite training sets
Nicolas L Roux, Mark Schmidt, and Francis R Bach · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
The tail at scale
Jeffrey Dean and Luiz André Barroso · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Solving the straggler problem with bounded staleness
James Cipar, Qirong Ho, Jin Kyu Kim, Seunghak Lee, Gregory R. Ganger, Garth Gibson, Kimberly Keeton, and Eric Xing · 2013
Earlier work this paper cites.
More effective distributed ml via a stale synchronous parallel parameter server
Qirong Ho, James Cipar, Henggang Cui, Seunghak Lee, Jin Kyu Kim, Phillip B Gibbons, Garth A Gibson, Greg Ganger, and Eric P Xing · 2013
Earlier work this paper cites.
Stochastic Processes: Theory for Applications
Robert G. Gallager · 2013
Earlier work this paper cites.
Efficient mini-batch training for stochastic optimization
Mu Li, Tong Zhang, Yuqiang Chen, and Alexander J. Smola · 2014
Earlier work this paper cites.
On the delay-storage trade-off in content download from coded distributed storage systems
Gauri Joshi, Yanpei Liu, and Emina Soljanin · 2014
Earlier work this paper cites.