Fetching the paper…
Reading the bibliography…
We provide the first theoretical analysis on the convergence rate of the asynchronous stochastic variance reduced gradient (SVRG) descent algorithm on non-convex optimization.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Learning multiple layers of features from tiny images, 2009
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
John Langford, Alexander Smola, and Martin Zinkevich · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
Distributed delayed stochastic optimization
Alekh Agarwal and John C Duchi · 2011
Earlier work this paper cites.
On optimization methods for deep learning
Jiquan Ngiam, Adam Coates, Ahbik Lahiri, Bobby Prochnow, Quoc V Le, and Andrew Y Ng · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al · 2012
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence _rate for finite training sets
Nicolas L Roux, Mark Schmidt, and Francis R Bach · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Cited alongside, same era.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Cited alongside, same era.
An asynchronous parallel randomized kaczmarz algorithm
Ji Liu, Stephen J Wright, and Srikrishna Sridhar · 2014
Cited alongside, same era.
Asynchronous distributed admm for consensus optimization
Ruiliang Zhang and James Kwok · 2014
Cited alongside, same era.
Deep learning
Fast distributed asynchronous sgd with variance reduction
Ruiliang Zhang, Shuai Zheng, and James T Kwok · 2015
Later among the works it cites.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
Martın Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al · 2016
Closest in time.
Variance reduction for faster non-convex optimization
Elad Allen-Zhu, Zeyuan; Hazan · 2016
Closest in time.
Asynchronous stochastic block coordinate descent with variance reduction
Bin Gu, Zhouyuan Huo, and Heng Huang · 2016
Closest in time.
Decoupled asynchronous proximal stochastic gradient descent with variance reduction
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
Asynchronous parallel stochastic gradient for nonconvex optimization
Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu · 2015
Cited alongside, same era.
On variance reduction in stochastic gradient descent and its asynchronous variants
Sashank J Reddi, Ahmed Hefny, Suvrit Sra, Barnabás Póczos, and Alex J Smola · 2015
Cited alongside, same era.
Zhouyuan Huo, Bin Gu, and Heng Huang · 2016
Closest in time.
Distributed asynchronous dual free stochastic dual coordinate ascent
Zhouyuan Huo and Heng Huang · 2016
Closest in time.
Stochastic variance reduction for nonconvex optimization
Sashank J Reddi, Ahmed Hefny, Suvrit Sra, Barnabás Póczós, and Alex Smola · 2016
Closest in time.
Fast asynchronous parallel stochastic gradient descent: A lock-free approach with convergence guarantee
Shen-Yi Zhao and Wu-Jun Li · 2016
Closest in time.