Fetching the paper…
Reading the bibliography…
Asynchronous momentum stochastic gradient descent algorithms (Async-MSGD) is one of the most popular algorithms in distributed machine learning.
A stochastic approximation method
Robbins, H · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Polyak, B. T · 1964
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E · 1986
Earlier work this paper cites.
Optimal unsupervised learning in a single-layer linear feedforward neural network
Sanger, T. D · 1989
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications, stochastic modelling and applied probability, vol. 35
Kushner, H. J · 2003
Earlier work this paper cites.
Restricted boltzmann machines for collaborative filtering
Salakhutdinov, R · 2007
Earlier work this paper cites.
Numerical methods for ordinary differential equations: initial value problems
Griffiths, D. F · 2010
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Hinton, G · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A · 2012
Earlier work this paper cites.
Weak convergence of probability measures
Sagitov, S · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P · 2014
Cited alongside, same era.
Scaling distributed machine learning with the parameter server
Li, M · 2014
Cited alongside, same era.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Ge, R · 2015
Cited alongside, same era.
Asynchronous parallel stochastic gradient for nonconvex optimization
Lian, X · 2015
Cited alongside, same era.
Staleness-aware async-sgd for distributed deep learning
Zhang, W · 2015
Symmetry, saddle points, and global geometry of nonconvex matrix factorization
Li, X · 2016
Later among the works it cites.
A comprehensive linear speedup analysis for asynchronous stochastic parallel optimization from zeroth-order to first-order
Lian, X · 2016
Later among the works it cites.
Asynchrony begets momentum, with an application to deep learning
Mitliagkas, I · 2016
Later among the works it cites.
A geometric analysis of phase retrieval
Sun, J · 2016
Later among the works it cites.
Online multiview representation learning: Dropping convexity for better efficiency
Chen, Z · 2017
Later among the works it cites.
Deep hyperspherical learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Revisiting distributed synchronous sgd
Chen, J · 2016
Cited alongside, same era.
Matrix completion has no spurious local minimum
Ge, R · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K · 2016
Cited alongside, same era.
Liu, W · 2017
Later among the works it cites.
Liu, T · 2018
Closest in time.
Yellowfin: Adaptive optimization for (a) synchronous systems
Zhang, J · 2018
Closest in time.