Fetching the paper…
Reading the bibliography…
We study convergence rates of AdaGrad-Norm as an exemplar of adaptive stochastic gradient methods (SGD), where the step sizes change based on observed stochastic gradients, for minimizing non-convex, smooth objectives.
“A Stochastic Approximation Method”
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
“Robust Regression and Lasso”
Huan Xu, Constantine Caramanis and Shie Mannor · 2008
Earlier work this paper cites.
“Measurement error models”
W.. Fuller · 2009
Earlier work this paper cites.
“Adaptive Bound Optimization for Online Convex Optimization”
H McMahan and Matthew Streeter · 2010
Earlier work this paper cites.
“Less Regret via Online Conditioning”
Matthew Streeter and H McMahan · 2010
Earlier work this paper cites.
“Adaptive Subgradient Methods for Online Learning and Stochastic Optimization”
John Duchi, Elad Hazan and Yoram Singer · 2011
Earlier work this paper cites.
“Stochastic First-and Zeroth-order Methods for Nonconvex Stochastic Programming”
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization”
Diederik. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
“Optimization Methods for Large-Scale Machine Learning”
Léon Bottou, Frank Curtis and Jorge Nocedal · 2018
Cited alongside, same era.
“On the Convergence of a Class of Adam-Type Algorithms for Non-Convex Optimization”
Xiangyi Chen, Sijia Liu, Ruoyu Sun and Mingyi Hong · 2018
Cited alongside, same era.
“Weighted AdaGrad with Unified Momentum”
Fangyu Zou, Li Shen, Zequn Jie, Ju Sun and Wei Liu · 2018
Cited alongside, same era.
“Lower Bounds for Non-Convex Stochastic Optimization”
Yossi Arjevani, Yair Carmon, John Duchi, Dylan Foster, Nathan Srebro and Blake Woodworth · 2019
Cited alongside, same era.
“Unixgrad: A universal, adaptive algorithm with optimal guarantees for constrained optimization”
Ali Kavis, Kfir Levy, Francis Bach and Volkan Cevher · 2019
“Asymptotic study of stochastic adaptive algorithm in non-convex landscape”
Sébastien Gadat and Ioana Gavra · 2020
Later among the works it cites.
“Feature Noise Induces Loss Discrepancy Across Groups”
Fereshte Khani and Percy Liang · 2020
Later among the works it cites.
“A High Probability Analysis of Adaptive SGD with Momentum”
Xiaoyu Li and Francesco Orabona · 2020
Later among the works it cites.
“On Stochastic Moving-Average Estimators for Non-Convex Optimization”
Zhishuai Guo, Yi Xu, Wotao Yin, Rong Jin and Tianbao Yang · 2021
Later among the works it cites.
“Domain-Independent Dominance of Adaptive Methods”
Pedro Savarese, David McAllester, Sudarshan Babu and Michael Maire · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“On the Convergence of Stochastic Gradient Descent with Adaptive Stepsizes”
Xiaoyu Li and Francesco Orabona · 2019
Cited alongside, same era.
“AdaGrad stepsizes: Sharp Convergence over Nonconvex Landscapes”
Rachel Ward, Xiaoxia Wu and Leon Bottou · 2019
Cited alongside, same era.
“On the Convergence of Adam and AdaGrad”
Alexandre Défossez, Léon Bottou, Francis Bach and Nicolas Usunier · 2020
Cited alongside, same era.
Ruinan Jin, Yu Xing and Xingkang He · 2022
Closest in time.
“High Probability Bounds for a Class of Nonconvex Algorithms with AdaGrad Stepsize”
Ali Kavis, Kfir Levy and Volkan Cevher · 2022
Closest in time.