Fetching the paper…
Reading the bibliography…
Asynchronous parallel implementations of stochastic gradient (SG) have been broadly used in solving deep neural network and received many successes in practice recently.
Parallel and distributed computation: numerical methods , volume 23
D. P. Bertsekas and J. N. Tsitsiklis · 1989
Earlier work this paper cites.
A neural probabilistic language model
Y. Bengio, R. Ducharme, P. Vincent, and C. Janvin · 2003
Earlier work this paper cites.
Online passive-aggressive algorithms
K. Crammer, O. Dekel, J. Keshet, S. Shalev-Shwartz, and Y. Singer · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Earlier work this paper cites.
Dual averaging method for regularized stochastic learning and online optimization
L. Xiao · 2009
Earlier work this paper cites.
Robust distributed online prediction
O. Dekel, R. Gilad-Bachrach, O. Shamir, and L. Xiao · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
M. Zinkevich, M. Weimer, L. Li, and A. J. Smola · 2010
Earlier work this paper cites.
Distributed delayed stochastic optimization
A. Agarwal and J. C. Duchi · 2011
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
E. Moulines and F. R. Bach · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
F. Niu, B. Recht, C. Re, and S. Wright · 2011
Earlier work this paper cites.
Online learning and online convex optimization
S. Shalev-Shwartz · 2011
Earlier work this paper cites.
Large scale distributed deep networks
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, A. Senior, P. Tucker, K. Yang, Q. V. Le, et al · 2012
Cited alongside, same era.
Optimal distributed online prediction using mini-batches
O. Dekel, R. Gilad-Bachrach, O. Shamir, and L. Xiao · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
Accelerated, parallel and proximal coordinate descent
O. Fercoq and P. Richtárik · 2013
Cited alongside, same era.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Cited alongside, same era.
Parameter server for distributed machine learning
M. Li, L. Zhou, Z. Yang, A. Li, F. Xia, D. G. Andersen, and A. Smola · 2013
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Later among the works it cites.
Asynchronous stochastic coordinate descent: Parallelism and convergence properties
J. Liu and S. J. Wright · 2014
Later among the works it cites.
Distributed block coordinate descent for minimizing partially separable functions
J. Marecek, P. Richtárik, and M. Takác · 2014
Later among the works it cites.
Gasgd: stochastic gradient descent for distributed asynchronous matrix completion via graph partitioning
F. Petroni and L. Querzoni · 2014
Later among the works it cites.
Regret bounded by gradual variation for online convex optimization
T. Yang, M. Mahdavi, R. Jin, and S. Zhu · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gpu asynchronous stochastic gradient descent to speed up neural network training
T. Paine, H. Jin, J. Yang, Z. Lin, and T. Huang · 2013
Cited alongside, same era.
An approximate, efficient LP solver for lp rounding
S. Sridhar, S. Wright, C. Re, J. Liu, V. Bittorf, and C. Zhang · 2013
Cited alongside, same era.
H. Yun, H.-F. Yu, C.-J. Hsieh, S. Vishwanathan, and I. Dhillon · 2013
Cited alongside, same era.
Revisiting asynchronous linear solvers: Provable convergence rate through randomization
H. Avron, A. Druinsky, and A. Gupta · 2014
Cited alongside, same era.
M. Hong · 2014
Cited alongside, same era.
Scaling distributed machine learning with the parameter server
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su
Cited in the paper.
Later among the works it cites.
Asynchronous distributed ADMM for consensus optimization
R. Zhang and J. Kwok · 2014
Later among the works it cites.
Deep learning with elastic averaging SGD
S. Zhang, A. Choromanska, and Y. LeCun · 2014
Later among the works it cites.
An asynchronous mini-batch algorithm for regularized stochastic optimization
H. R. Feyzmahdavian, A. Aytekin, and M. Johansson · 2015
Closest in time.
Perturbed iterate analysis for asynchronous stochastic optimization
H. Mania, X. Pan, D. Papailiopoulos, B. Recht, K. Ramchandran, and M. I. Jordan · 2015
Closest in time.
On the complexity of parallel coordinate descent
R. Tappenden, M. Takáč, and P. Richtárik · 2015
Closest in time.
Scaling up stochastic dual coordinate ascent
K. Tran, S. Hosseini, L. Xiao, T. Finley, and M. Bilenko · 2015
Closest in time.