Fetching the paper…
Reading the bibliography…
We prove the local convergence to minima and estimates on the rate of convergence for the stochastic gradient descent method in the case of not necessarily globally convex nor contracting objective functions.
Solving large scale linear prediction problems using stochastic gradient descent algorithms
T. Zhang · 1904
Earlier work this paper cites.
Regularity of the distance function
R. L. Foote · 1984
Earlier work this paper cites.
A stochastic approximation method
H. Robbins and S. Monro · 1985
Earlier work this paper cites.
Two problems with backpropagation and other steepest-descent learning procedures for networks
R. S. Sutton · 1986
Earlier work this paper cites.
Learning rate schedules for faster stochastic gradient search
C. Darken, J. Chang, and J. Moody · 1992
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Efficient backprop
Y. LeCun, L. Bottou, G. Orr, and K Muller · 1998
Earlier work this paper cites.
Natural gradient descent for on-line learning
M. Rattray, D. Saad, and S. I. Amari · 1998
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
N. Qian · 1999
Earlier work this paper cites.
On-line learning theory of soft committee machines with correlated hidden units steepest gradient descent and natural gradient descent
M. Inoue, H. Park, and M. Okada · 2003
Earlier work this paper cites.
Large scale online learning
L. Bottou and Y. LeCun · 2004
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
G. E. Hinton and R. R. Salakhutdinov · 2006
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
L. Bottou · 2010
Earlier work this paper cites.
An analysis on negative curvature induced by singularity in multi-layer neural-network learning
E. Mizutani and S. Dreyfus · 2010
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
F. Bach and E Moulines · 2011
Cited alongside, same era.
The tradeoffs of large scale learning
L. Bottou and O. Bousquet · 2011
Cited alongside, same era.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
E. Moulines and F. Bach · 2011
Cited alongside, same era.
Towards optimal one pass large scale learning with averaged stochastic gradient descent
W. Xu · 2011
Cited alongside, same era.
Large scale distributed deep networks
J. Dean, G. S. Corrado, R. Monga, K. Chen, M. Devin, Q. V. Le, M.C. Mao, M.A. Ranzato, A. Senior, P. Tucker, K. Yang, and A.Y. Ng · 2012
Cited alongside, same era.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Later among the works it cites.
Adaptivity of averaged stochastic gradient descent to local strong convexity for logistic regression
F. Bach · 2014
Later among the works it cites.
Revisiting natural gradient for deep networks
R. Pascanu and Y. Bengio · 2014
Later among the works it cites.
General multilevel adaptations for stochastic approximation algorithms
S. Dereich and T. Mueller-Gronbach · 2015
Later among the works it cites.
On the convergence rate of stochastic gradient descent for strongly convex functions
C. Tang and C. Monteleoni · 2015
Later among the works it cites.
Linear Convergence of Gradient and Proximal-Gradient Methods Under the Polyak-\LOjasiewicz Condition
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-R. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, and T.N. Sainath · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I Sutskever, and G. Hinton · 2012
Cited alongside, same era.
Making gradient descent optimal for strongly convex stochastic optimization
A. Rakhlin, O. Shamir, and K. Sridharan · 2012
Cited alongside, same era.
Generating sequences with recurrent neural networks
T. Schaul, S. Zhang, and Y. LeCun · 2012
Cited alongside, same era.
Non-strongly-convex smooth stochastic approximation with convergence rate o (1/n)
F. Bach and E. Moulines · 2013
Cited alongside, same era.
Generic stochastic gradient methods
B. Bercu and J.-C. Fort · 2013
Cited alongside, same era.
Recent advances in deep learning for speech research at Microsoft
L. Deng, J. Li, J.-T. Huang, K. Yao, D. Yu, F. Seide, M. Seltzer, G. Zweig, X. He, and J. Williams · 2013
Cited alongside, same era.
H. Karimi, J. Nutini, and M. Schmidt · 2016
Later among the works it cites.
An overview of gradient descent optimization algorithms
S. Ruder · 2016
Later among the works it cites.
Bridging the gap between constant step size stochastic gradient descent and Markov chains
A. Dieuleveut, A. Durmus, and B. Bach · 2017
Later among the works it cites.
Exponential convergence of testing error for stochastic gradient methods
L. Pillaud-Vivien, A. Rudi, and F. Bach · 2017
Later among the works it cites.
R. Vidal, J. Bruna, R. Giryes, and S. Soatto · 2017
Later among the works it cites.
Strong error analysis for stochastic gradient descent optimization algorithms
A. Jentzen, B. Kuckuck, A. Neufeld, and P. von Wurstemberger · 2018
Later among the works it cites.
On the convergence of stochastic gradient descent with adaptive stepsizes
X. Li and F. Orabona · 2018
Later among the works it cites.
Adagrad stepsizes: Sharp convergence over nonconvex landscapes, from any initialization
R. Ward, X. Wu, and L. Bottou · 2018
Later among the works it cites.