Fetching the paper…
Reading the bibliography…
Asynchronous stochastic gradient descent (ASGD) is a popular parallel optimization algorithm in machine learning.
On the theory of the brownian motion
George E Uhlenbeck and Leonard S Ornstein · 1930
Earlier work this paper cites.
Stochastic differential equations and applications
Xuerong Mao · 2007
Earlier work this paper cites.
Slow learners are fast
John Langford, Alex J Smola, and Martin Zinkevich · 2009
Earlier work this paper cites.
Probability: theory and examples
Rick Durrett · 2010
Earlier work this paper cites.
Distributed delayed stochastic optimization
Alekh Agarwal and John C Duchi · 2011
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al · 2012
Earlier work this paper cites.
Continuous-time stochastic mirror descent on a network: Variance reduction, consensus, convergence
Maxim Raginsky and J Bouvrie · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Cited alongside, same era.
Ergodicity for functional stochastic differential equations and applications
Jianhai Bao, George Yin, and Chenggui Yuan · 2014
Cited alongside, same era.
A differential equation for modeling nesterov’s accelerated gradient method: theory and insights
Weijie Su, Stephen Boyd, and Emmanuel Candes · 2014
Cited alongside, same era.
Accelerated mirror descent in continuous and discrete time
W Krichene, Am Bayen, and Pl Bartlett · 2015
Cited alongside, same era.
Asynchronous parallel stochastic gradient for nonconvex optimization
Online ica: Understanding global dynamics of nonconvex optimization via diffusion processes
Chris Junchi Li, Zhaoran Wang, and Han Liu · 2016
Later among the works it cites.
A variational analysis of stochastic gradient algorithms
Stephan Mandt, Matthew D Hoffman, and David M Blei · 2016
Later among the works it cites.
A variational perspective on accelerated methods in optimization
Andre Wibisono, Ashia C Wilson, and Michael I Jordan · 2016
Later among the works it cites.
Coupling adaptive batch sizes with learning rates
Lukas Balles, Javier Romero, and Philipp Hennig · 2017
Later among the works it cites.
Stochastic modified equations and adaptive stochastic gradient algorithms
Qianxiao Li, Cheng Tai, and Weinan E · 2017
Later among the works it cites.
The physical systems behind optimization algorithms
Lin F. Yang, R. Arora, V. Braverman, and Tuo Zhao · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu · 2015
Cited alongside, same era.
Deep learning with elastic averaging sgd
Sixin Zhang, Anna E Choromanska, and Yann LeCun · 2015
Cited alongside, same era.
Asymptotic analysis for functional stochastic differential equations, 2016
Jianhai Bao, George Yin, and Chenggui Yuan · 2016
Cited alongside, same era.
Later among the works it cites.
Asynchronous stochastic gradient descent with delay compensation
Shuxin Zheng, Qi Meng, Taifeng Wang, Wei Chen, Nenghai Yu, Zhi-Ming Ma, and Tie-Yan Liu · 2017
Later among the works it cites.