Fetching the paper…
Reading the bibliography…
In this paper, we propose a new, simplified high probability analysis of AdaGrad for smooth, non-convex problems.
A method for solving the convex programming problem with convergence rate o ( 1 / k 2 ) o(1/k^{2})
Yurii Nesterov · 1983
Earlier work this paper cites.
Introductory lectures on convex optimization. 2004, 2003
Yurii Nesterov · 2003
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich · 2003
Earlier work this paper cites.
On the generalization ability of online ¡i¿strongly¡/i¿ convex programming algorithms
Sham M. Kakade and Ambuj Tewari · 2008
Earlier work this paper cites.
Adaptive bound optimization for online convex optimization
H Brendan McMahan and Matthew Streeter · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Understanding Machine Learning: From Theory to Algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
Linear coupling: An ultimate unification of gradient and mirror descent, 2016
Zeyuan Allen-Zhu and Lorenzo Orecchia · 2016
Earlier work this paper cites.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
Saeed Ghadimi and Guanghui Lan · 2016
Earlier work this paper cites.
First-Order Methods in Optimization
Amir Beck · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Online adaptive methods, universality and acceleration
Kfir Y Levy, Alp Yurtsever, and Volkan Cevher · 2018
Cited alongside, same era.
On the convergence of adam and beyond
Sashank Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
Adaptive methods for nonconvex optimization
Manzil Zaheer, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
On the convergence of adaptive gradient methods for nonconvex optimization
Dongruo Zhou, Yiqi Tang, Ziyan Yang, Yuan Cao, and Quanquan Gu · 2018
AdaGrad stepsizes: Sharp convergence over nonconvex landscapes
Rachel Ward, Xiaoxia Wu, and Leon Bottou · 2019
Later among the works it cites.
A sufficient condition for convergences of adam and rmsprop
Fangyu Zou, Li Shen, Zequn Jie, Weizhong Zhang, and Wei Liu · 2019
Later among the works it cites.
Closing the generalization gap of adaptive gradient methods in training deep neural networks, 2020
Jinghui Chen, Dongruo Zhou, Yiqi Tang, Ziyan Yang, Yuan Cao, and Quanquan Gu · 2020
Later among the works it cites.
On the convergence of adam and adagrad, 03 2020
Alexandre Defossez, Leon Bottou, Francis Bach, and Nicolas Usunier · 2020
Later among the works it cites.
A simpler approach to accelerated optimization: iterative averaging meets optimism
Pooria Joulani, Anant Raj, Andras Gyorgy, and Csaba Szepesvari · 2020
Later among the works it cites.
First-order and Stochastic Optimization Methods for Machine Learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A general and adaptive robust loss function
Jonathan T. Barron · 2019
Cited alongside, same era.
On the convergence of a class of adam-type algorithms for non-convex optimization
Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong · 2019
Cited alongside, same era.
Nostalgic adam: Weighting more of the past gradients when designing the adaptive learning rate
Haiwen Huang, Chang Wang, and Bin Dong · 2019
Cited alongside, same era.
Unixgrad: A universal, adaptive algorithm with optimal guarantees for constrained optimization
Ali Kavis, Kfir Y. Levy, Francis Bach, and Volkan Cevher · 2019
Cited alongside, same era.
On the convergence of stochastic gradient descent with adaptive stepsizes
Xiaoyu Li and Francesco Orabona · 2019
Cited alongside, same era.
Adaptive gradient methods with dynamic bound of learning rate, 2019
Liangchen Luo, Yuanhao Xiong, Yan Liu, and Xu Sun · 2019
Cited alongside, same era.
Guanghui Lan · 2020
Later among the works it cites.
A high probability analysis of adaptive sgd with momentum, 2020
Xiaoyu Li and Francesco Orabona · 2020
Later among the works it cites.
An improved analysis of stochastic gradient descent with momentum, 2020
Yanli Liu, Yuan Gao, and Wotao Yin · 2020
Later among the works it cites.
Why are adaptive methods good for attention models?, 2020
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank J Reddi, Sanjiv Kumar, and Suvrit Sra · 2020
Later among the works it cites.
High-probability bounds for non-convex stochastic optimization with heavy tails, 2021
Ashok Cutkosky and Harsh Mehta · 2021
Later among the works it cites.
Momentum via primal averaging: Theoretical insights and learning rate schedules for non-convex optimization, 2021
Aaron Defazio · 2021
Later among the works it cites.
Adaptive gradient methods for constrained convex optimization and variational inequalities, 2021
Alina Ene, Huy L. Nguyen, and Adrian Vladu · 2021
Later among the works it cites.
STORM+: Fully adaptive SGD with recursive momentum for nonconvex optimization
Kfir Yehuda Levy, Ali Kavis, and Volkan Cevher · 2021
Later among the works it cites.