Fetching the paper…
Reading the bibliography…
We study the problem of parameter-free stochastic optimization, inquiring whether, and under what conditions, do fully parameter-free methods exist: these are methods that achieve convergence rates competitive with optimally tuned methods, without requiring significant knowledge of the true problem parameters.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
A. S. Nemirovskij and D. B. Yudin · 1983
Earlier work this paper cites.
A fast iterative shrinkage-thresholding algorithm for linear inverse problems
A. Beck and M. Teboulle · 2009
Earlier work this paper cites.
A parameter-free hedging algorithm
K. Chaudhuri, Y. Freund, and D. J. Hsu · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Stochastic gradient descent tricks
L. Bottou · 2012
Earlier work this paper cites.
An optimal method for stochastic composite optimization
G. Lan · 2012
Earlier work this paper cites.
No-regret algorithms for unconstrained online convex optimization
M. J. Streeter and H. B. McMahan · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Earlier work this paper cites.
No more pesky learning rates
T. Schaul, S. Zhang, and Y. LeCun · 2013
Earlier work this paper cites.
Unconstrained online linear learning in hilbert spaces: Minimax algorithms and normal approximations
H. B. McMahan and F. Orabona · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Achieving all with no parameters: Adanormalhedge
H. Luo and R. E. Schapire · 2015
Earlier work this paper cites.
Universal gradient methods for convex optimization problems
Y. Nesterov · 2015
Earlier work this paper cites.
Online convex optimization with unconstrained domains and losses
A. Cutkosky and K. A. Boahen · 2016
Earlier work this paper cites.
Coin betting and parameter-free online learning
F. Orabona and D. Pál · 2016
Cited alongside, same era.
Online learning without prior information
A. Cutkosky and K. Boahen · 2017
Cited alongside, same era.
Training deep networks without learning rates through coin betting
F. Orabona and T. Tommasi · 2017
Cited alongside, same era.
Black-box reductions for parameter-free online learning in banach spaces
A. Cutkosky and F. Orabona · 2018
Cited alongside, same era.
Scale-free online learning
F. Orabona and D. Pál · 2018
Cited alongside, same era.
On the convergence of adam and beyond
S. J. Reddi, S. Kale, and S. Kumar · 2018
Cited alongside, same era.
Artificial constraints and hints for unbounded online learning
A high probability analysis of adaptive sgd with momentum
X. Li and F. Orabona · 2020
Later among the works it cites.
Lipschitz and comparator-norm adaptivity in online learning
Z. Mhammedi and W. M. Koolen · 2020
Later among the works it cites.
Time-uniform, nonparametric, nonasymptotic confidence sequences
S. R. Howard, A. Ramdas, J. McAuliffe, and J. Sekhon · 2021
Later among the works it cites.
Parameter-free stochastic optimization of variationally coherent functions
F. Orabona and D. Pál · 2021
Later among the works it cites.
Lower bounds for non-convex stochastic optimization
Y. Arjevani, Y. Carmon, J. C. Duchi, D. J. Foster, N. Srebro, and B. Woodworth · 2022
Later among the works it cites.
Making sgd parameter-free
Y. Carmon and O. Hinder · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Cutkosky · 2019
Cited alongside, same era.
Revisiting the polyak step size
E. Hazan and S. Kakade · 2019
Cited alongside, same era.
Parameter-free online convex optimization with sub-exponential noise
K.-S. Jun and F. Orabona · 2019
Cited alongside, same era.
Unixgrad: A universal, adaptive algorithm with optimal guarantees for constrained optimization
A. Kavis, K. Y. Levy, F. Bach, and V. Cevher · 2019
Cited alongside, same era.
On the convergence of stochastic gradient descent with adaptive stepsizes
X. Li and F. Orabona · 2019
Cited alongside, same era.
On the convergence proof of amsgrad and a new version
P. T. Tran et al · 2019
Cited alongside, same era.
Later among the works it cites.
The power of adaptivity in sgd: Self-tuning step sizes with unbounded gradients and affine variance
M. Faw, I. Tziotis, C. Caramanis, A. Mokhtari, S. Shakkottai, and R. A. Ward · 2022
Later among the works it cites.
High probability bounds for a class of nonconvex algorithms with adagrad stepsize
A. Kavis, K. Y. Levy, and V. Cevher · 2022
Later among the works it cites.
Sgd with adagrad stepsizes: Full adaptivity with high probability to unknown parameters, unbounded gradients and affine variance
A. Attia and T. Koren · 2023
Later among the works it cites.
Learning-rate-free learning by d-adaptation
A. Defazio and K. Mishchenko · 2023
Later among the works it cites.
Dog is sgd’s best friend: A parameter-free dynamic step size schedule
M. Ivgi, O. Hinder, and Y. Carmon · 2023
Later among the works it cites.
High probability convergence of stochastic gradient methods
Z. Liu, T. D. Nguyen, T. H. Nguyen, A. Ene, and H. L. Nguyen · 2023
Later among the works it cites.
Prodigy: An expeditiously adaptive parameter-free learner
K. Mishchenko and A. Defazio · 2023
Later among the works it cites.
The price of adaptivity in stochastic convex optimization
Y. Carmon and O. Hinder · 2024
Closest in time.
Tuning-free stochastic optimization
A. Khaled and C. Jin · 2024
Closest in time.