Fetching the paper…
Reading the bibliography…
In this paper we propose several adaptive gradient methods for stochastic optimization.
A universal algorithm for variational inequalities adaptive to smoothness and noise
F. Bach and K. Y. Levy · 1902
Earlier work this paper cites.
Orth-method for smooth convex optimization
A. Nemirovski · 1982
Earlier work this paper cites.
Introduction to Optimization
B. Polyak · 1987
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Gradient methods for minimizing composite functions
Y. Nesterov · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images. phd thesis
A. Krizhevsky · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Sample size selection in optimization methods for machine learning
R. H. Byrd, G. M. Chin, J. Nocedal, and Y. Wu · 2012
Earlier work this paper cites.
Hybrid deterministic-stochastic methods for data fitting
M. P. Friedlander and M. Schmidt · 2012
Earlier work this paper cites.
Validation analysis of mirror descent stochastic approximation method
G. Lan, A. Nemirovski, and A. Shapiro · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Earlier work this paper cites.
First-order methods of smooth convex optimization with inexact oracle
O. Devolder, F. Glineur, and Y. Nesterov · 2014
Earlier work this paper cites.
Adam: a method for stochastic optimization
D. Kingma and J. Ba · 2015
Cited alongside, same era.
Universal gradient methods for convex optimization problems
Y. Nesterov · 2015
Cited alongside, same era.
A universal primal-dual convex optimization framework
A. Yurtsever, Q. Tran-Dinh, and V. Cevher · 2015
Cited alongside, same era.
Learning supervised pagerank with gradient-based and gradient-free optimization methods
L. Bogolubsky, P. Dvurechensky, A. Gasnikov, G. Gusev, Y. Nesterov, A. M. Raigorodskii, A. Tikhonov, and M. Zhukovskii · 2016
Cited alongside, same era.
Stochastic intermediate gradient method for convex optimization problems
A. V. Gasnikov and P. E. Dvurechensky · 2016
Cited alongside, same era.
Mathematical foundations of infinite-dimensional statistical models
A first-order primal-dual algorithm with linesearch
Y. Malitsky and T. Pock · 2018
Later among the works it cites.
Lectures on convex optimization
Y. Nesterov · 2018
Later among the works it cites.
Primal-dual accelerated gradient methods with small-dimensional relaxation oracle
Y. Nesterov, A. Gasnikov, S. Guminov, and P. Dvurechensky · 2018
Later among the works it cites.
Stochastic Gradient Descent: Recent Trends
D. Newton, F. Yousefian, and R. Pasupathy · 2018
Later among the works it cites.
Graph oracle models, lower bounds, and gaps for parallel stochastic optimization
B. E. Woodworth, J. Wang, A. Smith, B. McMahan, and N. Srebro · 2018
Later among the works it cites.
The complexity of finding stationary points with stochastic gradient descent
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Giné and R. Nickl · 2016
Cited alongside, same era.
An overview of gradient descent optimization algorithms
S. Ruder · 2016
Cited alongside, same era.
Lower bounds for finding stationary points ii: First-order methods
Y. Carmon, J. C. Duchi, O. Hinder, and A. Sidford · 2017
Cited alongside, same era.
Extragradient method with variance reduction for stochastic variational inequalities
A. N. Iusem, A. Jofré, R. I. Oliveira, and P. Thompson · 2017
Cited alongside, same era.
Optimal adaptive and accelerated stochastic gradient descent
Q. Deng, Y. Cheng, and G. Lan · 2018
Cited alongside, same era.
P. Dvurechensky, A. Gasnikov, and A. Kroshnin · 2018
Cited alongside, same era.
About the power law of the pagerank vector distribution. part 2. backley–osthus model, power law verification for this model and setup of real search engines
A. V. Gasnikov, P. Dvurechenskii, M. E. Zhukovskii, S. V. Kim, S. S. Plaunov, D. A. Smirnov, and F. A. Noskov · 2018
Cited alongside, same era.
Y. Drori and O. Shamir · 2019
Closest in time.
On dual approach for distributed stochastic convex optimization over networks
D. Dvinskikh, E. Gorbunov, A. Gasnikov, P. Dvurechensky, and C. A. Uribe · 2019
Closest in time.
Optimal mini-batch and step sizes for saga
N. Gazagnadou, R. M. Gower, and J. Salmon · 2019
Closest in time.
Optimal decentralized distributed algorithms for stochastic convex optimization
E. Gorbunov, D. Dvinskikh, and A. Gasnikov · 2019
Closest in time.
Variance-based extragradient methods with line search for stochastic variational inequalities
A. N. Iusem, A. Jofré, R. I. Oliveira, and P. Thompson · 2019
Closest in time.
Unixgrad: A universal, adaptive algorithm with optimal guarantees for constrained optimization
A. Kavis, K. Y. Levy, F. Bach, and V. Cevher · 2019
Closest in time.
Heuristic adaptive fast gradient method in stochastic optimization tasks
A. Ogaltsov and A. Tyurin · 2019
Closest in time.
AdaGrad stepsizes: Sharp convergence over nonconvex landscapes
R. Ward, X. Wu, and L. Bottou · 2019
Closest in time.