Fetching the paper…
Reading the bibliography…
We study the problem of stochastic optimization for deep learning in the parallel computing environment under communication constraints.
Optimization theory: the finite dimensional case
Hestenes, M. R · 1975
Earlier work this paper cites.
Parallel and Distributed Computation
Bertsekas, D. P and Tsitsiklis, J. N · 1989
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Polyak, B. T and Juditsky, A. B · 1992
Earlier work this paper cites.
Online algorithms and stochastic approximations
Bottou, L · 1998
Earlier work this paper cites.
Asynchronous stochastic approximations
Borkar, V · 1998
Earlier work this paper cites.
Distributed asynchronous incremental subgradient methods
Nedić, A, Bertsekas, D, and Borkar, V · 2001
Earlier work this paper cites.
On the generalization ability of on-line learning algorithms
Cesa-Bianchi, N, Conconi, A, and Gentile, C · 2004
Earlier work this paper cites.
Introductory lectures on convex optimization
Nesterov, Y · 2004
Earlier work this paper cites.
Smooth minimization of non-smooth functions
Nesterov, Y · 2005
Earlier work this paper cites.
Numerical Optimization, Second Edition
Nocedal, J and Wright, S · 2006
Earlier work this paper cites.
Slow learners are fast
Langford, J, Smola, A, and Zinkevich, M · 2009
Earlier work this paper cites.
Parallelized stochastic gradient descent
Zinkevich, M, Weimer, M, Smola, A, and Li, L · 2010
Cited alongside, same era.
Distributed optimization and statistical learning via the alternating direction method of multipliers
Boyd, S, Parikh, N, Chu, E, Peleato, B, and Eckstein, J · 2011
Cited alongside, same era.
Scaling up machine learning: Parallel and distributed approaches
Bekkerman, R, Bilenko, M, and Langford, J · 2011
Cited alongside, same era.
Distributed delayed stochastic optimization
Agarwal, A and Duchi, J · 2011
Cited alongside, same era.
Hogwild: A Lock-Free Approach to Parallelizing Stochastic Gradient Descent
Recht, B, Re, C, Wright, S. J, and Niu, F · 2011
Cited alongside, same era.
Large scale distributed deep networks
Dean, J, Corrado, G, Monga, R, Chen, K, Devin, M, Le, Q, Mao, M, Ranzato, M, Senior, A, Tucker, P, Yang, K, and Ng, A · 2012
Gpu asynchronous stochastic gradient descent to speed up neural network training
Paine, T, Jin, H, Yang, J, Lin, Z, and Huang, T · 2013
Later among the works it cites.
More effective distributed ml via a stale synchronous parallel parameter server
Ho, Q, Cipar, J, Cui, H, Lee, S, Kim, J. K, Gibbons, P. B, Gibson, G. A, Ganger, G, and Xing, E. P · 2013
Later among the works it cites.
On the importance of initialization and momentum in deep learning
Sutskever, I, Martens, J, Dahl, G, and Hinton, G · 2013
Later among the works it cites.
Stochastic alternating direction method of multipliers
Ouyang, H, He, N, Tran, L, and Gray, A · 2013
Later among the works it cites.
Regularization of neural networks using dropconnect
Wan, L, Zeiler, M. D, Zhang, S, LeCun, Y, and Fergus, R · 2013
Later among the works it cites.
Fundamental limits of online and distributed algorithms for statistical learning and estimation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A, Sutskever, I, and Hinton, G. E · 2012
Cited alongside, same era.
An optimal method for stochastic composite optimization
Lan, G · 2012
Cited alongside, same era.
OverFeat: Integrated Recognition, Localization and Detection using Convolutional Networks
Sermanet, P, Eigen, D, Zhang, X, Mathieu, M, Fergus, R, and LeCun, Y · 2013
Cited alongside, same era.
Multi-gpu training of convnets
Yadan, O, Adams, K, Taigman, Y, and Ranzato, M · 2013
Cited alongside, same era.
Shamir, O · 2014
Closest in time.
1-bit stochastic gradient descent and application to data-parallel distributed training of speech dnns
Seide, F, Fu, H, Droppo, J, Li, G, and Yu, D · 2014
Closest in time.
Towards an optimal stochastic alternating direction method of multipliers
Azadi, S and Sra, S · 2014
Closest in time.
Asynchronous distributed admm for consensus optimization
Zhang, R and Kwok, J · 2014
Closest in time.
The loss surfaces of multilayer networks
Choromanska, A, Henaff, M. B, Mathieu, M, Arous, G. B, and LeCun, Y · 2015
Closest in time.