Fetching the paper…
Reading the bibliography…
In this work, we describe a generic approach to show convergence with high probability for both stochastic convex and non-convex optimization with sub-Gaussian noise.
On tail probabilities for martingales
David A Freedman · 1975
Earlier work this paper cites.
Concentration inequalities and martingale inequalities: a survey
Fan Chung and Linyuan Lu · 2006
Earlier work this paper cites.
On the generalization ability of online strongly convex programming algorithms
Sham M Kakade and Ambuj Tewari · 2008
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Solving variational inequalities with stochastic mirror-prox algorithm
Anatoli Juditsky, Arkadi Nemirovski, and Claire Tauvel · 2011
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2011
Earlier work this paper cites.
An optimal method for stochastic composite optimization
Guanghui Lan · 2012
Earlier work this paper cites.
Beyond the regret minimization barrier: optimal algorithms for stochastic strongly-convex optimization
Elad Hazan and Satyen Kale · 2014
Earlier work this paper cites.
Stochastic intermediate gradient method for convex problems with stochastic inexact oracle
Pavel Dvurechensky and Alexander Gasnikov · 2016
Earlier work this paper cites.
High-dimensional probability: An introduction with applications in data science , volume 47
Roman Vershynin · 2018
Cited alongside, same era.
Tight analyses for non-smooth stochastic gradient descent
Nicholas JA Harvey, Christopher Liaw, Yaniv Plan, and Sikander Randhawa · 2019
Cited alongside, same era.
On the convergence of stochastic gradient descent with adaptive stepsizes
Xiaoyu Li and Francesco Orabona · 2019
Cited alongside, same era.
Algorithms of robust stochastic optimization based on mirror descent method
Alexander V Nazin, Arkadi S Nemirovsky, Alexandre B Tsybakov, and Anatoli B Juditsky · 2019
Cited alongside, same era.
Adagrad stepsizes: Sharp convergence over nonconvex landscapes
Rachel Ward, Xiaoxia Wu, and Leon Bottou · 2019
Cited alongside, same era.
Stochastic optimization with heavy-tailed noise via accelerated gradient clipping
Why are adaptive methods good for attention models?
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank Reddi, Sanjiv Kumar, and Suvrit Sra · 2020
Later among the works it cites.
High-probability bounds for non-convex stochastic optimization with heavy tails
Ashok Cutkosky and Harsh Mehta · 2021
Later among the works it cites.
From low probability to high confidence in stochastic convex optimization
Damek Davis, Dmitriy Drusvyatskiy, Lin Xiao, and Junyu Zhang · 2021
Later among the works it cites.
Adaptive gradient methods for constrained convex optimization and variational inequalities
Alina Ene, Huy L Nguyen, and Adrian Vladu · 2021
Later among the works it cites.
High probability bounds for a class of nonconvex algorithms with adagrad stepsize
Ali Kavis, Kfir Yehuda Levy, and Volkan Cevher · 2021
Later among the works it cites.
A simple convergence proof of adam and adagrad
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Eduard Gorbunov, Marina Danilova, and Alexander Gasnikov · 2020
Cited alongside, same era.
First-order and stochastic optimization methods for machine learning
Guanghui Lan · 2020
Cited alongside, same era.
A high probability analysis of adaptive sgd with momentum
Xiaoyu Li and Francesco Orabona · 2020
Cited alongside, same era.
High probability convergence and uniform stability bounds for nonconvex stochastic gradient descent
Liam Madden, Emiliano Dall’Anese, and Stephen Becker · 2020
Cited alongside, same era.
Alexandre Défossez, Léon Bottou, Francis Bach, and Nicolas Usunier · 2022
Later among the works it cites.
The power of adaptivity in sgd: Self-tuning step sizes with unbounded gradients and affine variance
Matthew Faw, Isidoros Tziotis, Constantine Caramanis, Aryan Mokhtari, Sanjay Shakkottai, and Rachel Ward · 2022
Later among the works it cites.
High probability guarantees for nonconvex stochastic gradient descent with heavy tails
Shaojie Li and Yong Liu · 2022
Later among the works it cites.
Zijian Liu, Ta Duy Nguyen, Alina Ene, and Huy L Nguyen · 2022
Later among the works it cites.