Fetching the paper…
Reading the bibliography…
This paper proposes a new easy-to-implement parameter-free gradient-based optimizer: DoWG (Distance over Weighted Gradients).
Revisiting the Polyak step size
Elad Hazan and Sham Kakade · 1905
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Boris T. Polyak · 1964
Earlier work this paper cites.
Introduction to optimization
Boris Polyak · 1987
Earlier work this paper cites.
How to use expert advice
Nicolò Cesa-Bianchi, Yoav Freund, David Haussler, David P. Helmbold, Robert E. Schapire, and Manfred K. Warmuth · 1997
Earlier work this paper cites.
Better theory for SGD in the nonconvex world
Ahmed Khaled and Peter Richtárik · 2002
Earlier work this paper cites.
Convex Optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
The cost of training NLP models: a concise overview
Or Sharir, Barak Peleg, and Yoav Shoham · 2004
Earlier work this paper cites.
Prediction, Learning, and Games
Nicolo Cesa-Bianchi and Gabor Lugosi · 2006
Earlier work this paper cites.
Online learning with prior knowledge
Elad Hazan and Nimrod Megiddo · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John C. Duchi, Elad Hazan, and Yoram Singer · 2010
Earlier work this paper cites.
LIBSVM: A library for support vector machines
Chih-Chung Chang and Chih-Jen Lin · 2011
Earlier work this paper cites.
Minimization methods for non-differentiable functions , volume 3
Naum Zuselevich Shor · 2012
Earlier work this paper cites.
No-regret algorithms for unconstrained online convex optimization
Matthew Streeter and H. Brendan McMahan · 2012
Earlier work this paper cites.
Dimension-free exponentiated gradient
Francesco Orabona · 2013
Earlier work this paper cites.
Unconstrained online linear learning in hilbert spaces: Minimax algorithms and normal approximations
H. Brendan McMahan and Francesco Orabona · 2014
Earlier work this paper cites.
Universal gradient methods for convex optimization problems
Yurii Nesterov · 2014
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Sébastien Bubeck · 2015
Earlier work this paper cites.
Beyond convexity: Stochastic quasi-convex optimization
Elad Hazan, Kfir Y. Levy, and Shai Shalev-Shwartz · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Practical Methodology , chapter 11
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Coin betting and parameter-free online learning
Francesco Orabona and Dávid Pál · 2016
Cited alongside, same era.
A unified approach to adaptive regularization in online and stochastic optimization
Vineet Gupta, Tomer Koren, and Yoram Singer · 2017
Cited alongside, same era.
Online to offline conversions, universality and adaptive minibatch sizes
Kfir Y. Levy · 2017
Cited alongside, same era.
Training deep networks without learning rates through coin betting
Francesco Orabona and Tatiana Tommasi · 2017
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
Ashia C. Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Why are adaptive methods good for attention models?
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank J. Reddi, Sanjiv Kumar, and Suvrit Sra · 2020
Later among the works it cites.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy M. Cohen, Simran Kaur, Yuanzhi Li, J. Zico Kolter, and Ameet Talwalkar · 2021
Later among the works it cites.
Adaptive gradient methods for constrained convex optimization and variational inequalities
Alina Ene, Huy L. Nguyen, and Adrian Vladu · 2021
Later among the works it cites.
Stochastic polyak step-size for SGD: an adaptive learning rate for fast convergence
Nicolas Loizou, Sharan Vaswani, Issam Hadj Laradji, and Simon Lacoste-Julien · 2021
Later among the works it cites.
Parameter-free stochastic optimization of variationally coherent functions
Francesco Orabona and Dávid Pál · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E. Curtis, and Jorge Nocedal · 2018
Cited alongside, same era.
Online adaptive methods, universality and acceleration
Kfir Yehuda Levy, Alp Yurtsever, and Volkan Cevher · 2018
Cited alongside, same era.
Lectures on Convex Optimization
Yurii Nesterov · 2018
Cited alongside, same era.
Artificial constraints and hints for unbounded online learning
Ashok Cutkosky · 2019
Cited alongside, same era.
Convergence rates for deterministic and stochastic subgradient methods without lipschitz continuity
Benjamin Grimmer · 2019
Cited alongside, same era.
UniXGrad: A universal, adaptive algorithm with optimal guarantees for constrained optimization
Ali Kavis, Kfir Y. Levy, Francis R. Bach, and Volkan Cevher · 2019
Cited alongside, same era.
David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean · 2021
Later among the works it cites.
Sequential convergence of AdaGrad algorithm for smooth convex optimization
Cheik Traoré and Edouard Pauwels · 2021
Later among the works it cites.
Understanding gradient descent on edge of stability in deep learning
Sanjeev Arora, Zhiyuan Li, and Abhishek Panigrahi · 2022
Later among the works it cites.
Making SGD parameter-free
Yair Carmon and Oliver Hinder · 2022
Later among the works it cites.
Adaptive gradient methods at the edge of stability
Jeremy M. Cohen, Behrooz Ghorbani, Shankar Krishnan, Naman Agarwal, Sourabh Medapati, Michal Badura, Daniel Suo, David Cardoze, Zachary Nado, George E. Dahl, and Justin Gilmer · 2022
Later among the works it cites.
On optimal universal first-order methods for minimizing heterogeneous sums
Benjamin Grimmer · 2022
Later among the works it cites.
Zijian Liu, Ta Duy Nguyen, Alina Ene, and Huy L. Nguyen · 2022
Later among the works it cites.
Antonio Orvieto, Simon Lacoste-Julien, and Nicolas Loizou · 2022
Later among the works it cites.
Toward understanding why Adam converges faster than SGD for transformers
Yan Pan and Yuanzhi Li · 2022
Later among the works it cites.
Self-stabilization: The implicit bias of gradient descent at the edge of stability
Alex Damian, Eshaan Nichani, and Jason D. Lee · 2023
Closest in time.
Learning-rate-free learning by D-adaptation
Aaron Defazio and Konstantin Mishchenko · 2023
Closest in time.
DoG is SGD’s best friend: A parameter-free dynamic step size schedule
Maor Ivgi, Oliver Hinder, and Yair Carmon · 2023
Closest in time.
Noise is not the main factor behind the gap between SGD and Adam on transformers, but sign descent might be
Frederik Kunstner, Jacques Chen, Jonathan Wilder Lavington, and Mark Schmidt · 2023
Closest in time.
Puya Latafat, Andreas Themelis, Lorenzo Stella, and Panagiotis Patrinos · 2023
Closest in time.
Convergence of adam under relaxed assumptions
Haochuan Li, Ali Jadbabaie, and Alexander Rakhlin · 2023
Closest in time.
Francesco Orabona · 2023
Closest in time.