Fetching the paper…
Reading the bibliography…
This paper studies some asymptotic properties of adaptive algorithms widely used in optimization and machine learning, and among them Adagrad and Rmsprop, which are involved in most of the blackbox deep learning algorithms.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. Polyak · 1964
Earlier work this paper cites.
A convergence theorem for non negative almost supermartingales and some applications
H. Robbins and D. Siegmund · 1971
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o (1/kˆ 2)
Y. E Nesterov · 1983
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
A. S. Nemirovski and D. B. Yudin · 1983
Earlier work this paper cites.
Théorèmes de convergence presque sure pour une classe d’algorithmes stochastiques à pas décroissant
M. Métivier and P. Priouret · 1987
Earlier work this paper cites.
Efficient estimations from a slowly convergent robbins-monro process
D. Ruppert · 1988
Earlier work this paper cites.
Nonconvergence to unstable points in urn models and stochastic approximations
R. Pemantle · 1990
Earlier work this paper cites.
Systemes dynamiques dissipatifs et applications
A. Haraux · 1991
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. Juditsky · 1992
Earlier work this paper cites.
Dynamics of morse-smale urn processes
M. Benaïm and Morris W Hirsch · 1995
Earlier work this paper cites.
Les algorithmes stochastiques contournent-ils les pieges?
O. Brandiere and M. Duflo · 1996
Earlier work this paper cites.
Asymptotic pseudotrajectories and chain recurrent flows, with applications
M. Benaïm and M. Hirsch · 1996
Earlier work this paper cites.
Stochastic approximation with two time scales
V.S. Borkar · 1997
Earlier work this paper cites.
Dynamics of stochastic approximation algorithms
M. Benaïm · 1999
Earlier work this paper cites.
The heavy ball with friction method, i. the continuous dynamical system
H. Attouch, X. Goudou, and P. Redont · 2000
Earlier work this paper cites.
A second-order gradient-like dissipative dynamical system with hessian-driven damping.: Application to optimization and mechanics
F. Alvarez, H. Attouch, J. Bolte, and P. Redont · 2002
Earlier work this paper cites.
Introductory lectures on convex optimization
Y. Nesterov · 2004
Earlier work this paper cites.
Convergence rate and averaging of nonlinear two-time-scale stochastic approximation algorithms
A. Mokkadem and M. Pelletier · 2006
Earlier work this paper cites.
The tradeoffs of large scale learning
L. Bottou and O. Bousquet · 2007
Cited alongside, same era.
On the long time behavior of second order differential equations with asymptotically small dissipation
A. Cabot, H. Engler, and S. Gadat · 2009
Cited alongside, same era.
Second-order differential equations with asymptotically small dissipation and piecewise flat potentials
A. Cabot, H. Engler, and S. Gadat · 2009
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
E. Moulines and F. Bach · 2011
Cited alongside, same era.
A stochastic gradient method with an exponential convergence rate for finite training sets
Optimization methods for large-scale machine learning
L. Bottou, F. E Curtis, and J. Nocedal · 2018
Later among the works it cites.
Stochastic heavy ball
S. Gadat, F. Panloup, and S. Saadane · 2018
Later among the works it cites.
Accelerated gradient descent escapes saddle points faster than gradient descent
C. Jin, P. Netrapalli, and M.I. Jordan · 2018
Later among the works it cites.
Rate of convergence of the Nesterov accelerated gradient method in the subcritical case 3
H. Attouch, Z. Chbani, and H. Riahi · 2019
Later among the works it cites.
AdaGrad stepsizes: Sharp convergence over nonconvex landscapes
R. Ward, X. Wu, and L. Bottou · 2019
Later among the works it cites.
A sufficient condition for convergences of Adam and RMSProp
F. Zou, F. Shen, Z. Jie, W. Zhang, and W. Liu · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Le Roux, M. Schmidt, F. R Bach, et al · 2012
Cited alongside, same era.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Cited alongside, same era.
Adaptivity of averaged stochastic gradient descent to local strong convexity for logistic regression
F. Bach · 2014
Cited alongside, same era.
Long time behaviour and stationary regime of memory gradient diffusions
S. Gadat and F. Panloup · 2014
Cited alongside, same era.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Later among the works it cites.
Convergence and dynamical behavior of the Adam algorithm for non-convex stochastic optimization
A. Barakat and P. Bianchi · 2020
Closest in time.
Statistical properties of stochastic newton algorithms for regularized semi-discrete optimal transport, Preprint
B. Bercu, J. Bigot, S. Gadat, and E. Siviero · 2020
Closest in time.
Stochastic optimization with momentum: convergence, fluctuations, and traps avoidance
A. Barakat, P. Bianchi, W. Hachem, and Sh. Schechtman · 2020
Closest in time.
Stochastic estimation algorithm for superquantile estimation
B. Bercu, M. Costa, and S. Gadat · 2020
Closest in time.
A general system of differential equations to model first-order adaptive algorithms
A. Belotto da Silva and M. Gazeau · 2020
Closest in time.
An efficient stochastic newton algorithm for parameter estimation in logistic regressions
B. Bercu, A. Godichon, and B. Portier · 2020
Closest in time.
Non-asymptotic study of a recursive superquantile estimation algorithm
M. Costa and S. Gadat · 2020
Closest in time.
An efficient averaged stochastic gauss-newton algorithm for estimating parameters of non linear regressions models, 2020
P. Cénac, A. Godichon-Baggioni, and B. Portier · 2020
Closest in time.
On the convergence of Adam and Adagrad
A. Défossez, L. Bottou, F. Bach, and N. Usunier · 2020
Closest in time.
Optimal non-asymptotic bound of the Ruppert-Polyak averaging without strong convexity
S. Gadat and F. Panloup · 2020
Closest in time.
Momentum and stochastic momentum for stochastic gradient, newton, proximal point and subspace descent methods
N. Loizou and P. Richtárik · 2020
Closest in time.
On the convergence of the stochastic heavy ball method
O. Sebbouh, R. M Gower, and A. Defazio · 2020
Closest in time.