Fetching the paper…
Reading the bibliography…
Deep learning optimizers are often motivated through a mix of convex and approximate second-order theory.
Minimization of functions having Lipschitz continuous first partial derivatives
Larry Armijo · 1966
Earlier work this paper cites.
Some iterative methods for improving orthonormality
Zdislav Kovarik · 1970
Earlier work this paper cites.
An iterative algorithm for computing the best estimate of an orthogonal matrix
Åke Björck and C. Bowie · 1971
Earlier work this paper cites.
A direct adaptive method for faster backpropagation learning: The RPROP algorithm
Martin Riedmiller and Heinrich Braun · 1993
Earlier work this paper cites.
On the computation of the matrix k-th root
Slobodan Lakić · 1998
Earlier work this paper cites.
Functions of Matrices
Nicholas J. Higham · 2008
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John C. Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Stochastic spectral descent for discrete graphical models
David Carlson, Ya-Ping Hsieh, Edo Collins, Lawrence Carin, and Volkan Cevher · 2016
Earlier work this paper cites.
MM Optimization Algorithms
Kenneth Lange · 2016
Earlier work this paper cites.
Kai Fan · 2017
Earlier work this paper cites.
A unified approach to adaptive regularization in online and stochastic optimization
Vineet Gupta, Tomer Koren, and Yoram Singer · 2017
Earlier work this paper cites.
Dissecting Adam: The sign, magnitude and variance of stochastic gradients
Lukas Balles and Philipp Hennig · 2018
Cited alongside, same era.
signSGD: Compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar · 2018
Cited alongside, same era.
Shampoo: Preconditioned stochastic tensor optimization
Vineet Gupta, Tomer Koren, and Yoram Singer · 2018
Cited alongside, same era.
Scalable second order optimization for deep learning
Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, and Yoram Singer · 2020
Cited alongside, same era.
Randomized numerical linear algebra: Foundations and algorithms
Per-Gunnar Martinsson and Joel A. Tropp · 2020
Cited alongside, same era.
Connection of diagonal Hessian estimates to natural gradients in stochastic optimization
DoG is SGD’s best friend: A parameter-free dynamic step size schedule
Maor Ivgi, Oliver Hinder, and Yair Carmon · 2023
Later among the works it cites.
DoWG unleashed: An efficient universal parameter-free gradient descent method
Ahmed Khaled, Konstantin Mishchenko, and Chi Jin · 2023
Later among the works it cites.
Prodigy: An expeditiously adaptive parameter-free learner
Konstantin Mishchenko and Aaron Defazio · 2023
Later among the works it cites.
Hao-Jun Michael Shi, Tsung-Hsien Lee, Shintaro Iwasaki, Jose Gallego-Posada, Zhijing Li, Kaushik Rangadurai, Dheevatsa Mudigere, and Michael Rabbat · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shiqing Sun and James C. Spall · 2021
Cited alongside, same era.
Automatic Gradient Descent: Deep Learning without Hyperparameters
Jeremy Bernstein, Chris Mingard, Kevin Huang, Navid Azizan, and Yisong Yue · 2023
Cited alongside, same era.
Symbolic discovery of optimization algorithms
Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, Yifeng Lu, and Quoc V Le · 2023
Cited alongside, same era.
Benchmarking neural network training algorithms
George E. Dahl, Frank Schneider, Zachary Nado, Naman Agarwal, Chandramouli Shama Sastry, Philipp Hennig, Sourabh Medapati, Runa Eschenhagen, Priya Kasimbeg, Daniel Suo, Juhan Bae, Justin Gilmer, Abel L. Peirson, Bilal Khan, Rohan Anil, Mike Rabbat, Shankar Krishnan, Daniel Snider, Ehsan Amid, Kongtao Chen, Chris J. Maddison, Rakshith Vasudev, Michal Badura, Ankush Garg, and Peter Mattson · 2023
Cited alongside, same era.
Learning-rate-free learning by D-adaptation
Aaron Defazio and Konstantin Mishchenko · 2023
Cited alongside, same era.
Sketchy: Memory-efficient adaptive regularization with frequent directions
Vladimir Feinberg, Xinyi Chen, Y. Jennifer Sun, Rohan Anil, and Elad Hazan · 2023
Cited alongside, same era.
Stochastic spectral descent for Restricted Boltzmann Machines
David Carlson, Volkan Cevher, and Lawrence Carin
Cited in the paper.
Matthew Streeter · 2023
Later among the works it cites.
A spectral condition for feature learning
Greg Yang, James B. Simon, and Jeremy Bernstein · 2023
Later among the works it cites.
Improving line search methods for large scale neural network training
Philip Kenneweg, Tristan Kenneweg, and Barbara Hammer · 2024
Closest in time.
Scalable optimization in the modular norm
Tim Large, Yang Liu, Minyoung Huh, Hyojin Bahng, Phillip Isola, and Jeremy Bernstein · 2024
Closest in time.
A new perspective on Shampoo’s preconditioner
Depen Morwani, Itai Shapira, Nikhil Vyas, Eran Malach, Sham Kakade, and Lucas Janson · 2024
Closest in time.
Implicit bias of AdamW: ℓ ∞ \ell_{\infty} -norm constrained optimization
Shuo Xie and Zhiyuan Li · 2024
Closest in time.
Deconstructing what makes a good optimizer for language models
Rosie Zhao, Depen Morwani, David Brandfonbrener, Nikhil Vyas, and Sham Kakade · 2024
Closest in time.