Fetching the paper…
Reading the bibliography…
Lion (Evolved Sign Momentum), a new optimizer discovered through program search, has shown promising results in training large AI models.
Revisiting the Polyak step size, August 2022
Elad Hazan and Sham Kakade · 1905
Earlier work this paper cites.
An algorithm for quadratic programming
M. Frank and P. Wolfe · 1956
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Arkadij Semenovic Nemirovskij and David Borisovich Yudin · 1983
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o (1/ κ \kappa ˆ 2)
Yurii Evgen’evich Nesterov · 1983
Earlier work this paper cites.
Convex Analysis , volume 11
R. T. Rockafellar · 1997
Earlier work this paper cites.
Conformal hamiltonian systems
Robert McLachlan and Matthew Perlmutter · 2001
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman, Geoffrey Hinton, et al · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Accelerated mirror descent in continuous and discrete time
Walid Krichene, Alexandre Bayen, and Peter L Bartlett · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Cited alongside, same era.
Fast convex optimization via inertial dynamics with hessian driven damping
Hedy Attouch, Juan Peypouquet, and Patrick Redont · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
The power of normalization: Faster evasion of saddle points
Kfir Y Levy · 2016
Cited alongside, same era.
Neural optimizer search with reinforcement learning
Revisiting normalized gradient descent: Fast evasion of saddle points
Ryan Murray, Brian Swenson, and Soummya Kar · 2019
Later among the works it cites.
Hamiltonian descent for composite objectives
Brendan O’Donoghue and Chris J Maddison · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
Pyglove: Symbolic programming for automated machine learning
Daiyi Peng, Xuanyi Dong, Esteban Real, Mingxing Tan, Yifeng Lu, Gabriel Bender, Hanxiao Liu, Adam Kraft, Chen Liang, and Quoc Le · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Irwan Bello, Barret Zoph, Vijay Vasudevan, and Quoc V Le · 2017
Cited alongside, same era.
Adam: A method for stochastic optimization, 2017
Diederik P. Kingma and Jimmy Ba · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Dissecting adam: The sign, magnitude and variance of stochastic gradients
Lukas Balles and Philipp Hennig · 2018
Cited alongside, same era.
Chris J Maddison, Daniel Paulin, Yee Whye Teh, Brendan O’Donoghue, and Arnaud Doucet · 2018
Cited alongside, same era.
Decoupled weight decay regularization, 2019
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
signSGD: Compressed Optimisation for Non-Convex Problems, August 2018a
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Anima Anandkumar
Cited in the paper.
Automl-zero: Evolving machine learning algorithms from scratch
Esteban Real, Chen Liang, David So, and Quoc Le · 2020
Later among the works it cites.
Dual space preconditioning for gradient descent
Chris J Maddison, Daniel Paulin, Yee Whye Teh, and Arnaud Doucet · 2021
Later among the works it cites.
Understanding the acceleration phenomenon via high-resolution differential equations
Bin Shi, Simon S Du, Michael I Jordan, and Weijie J Su · 2021
Later among the works it cites.
Robustness to unbounded smoothness of generalized signsgd
Michael Crawshaw, Mingrui Liu, Francesco Orabona, Wei Zhang, and Zhenxun Zhuang · 2022
Later among the works it cites.
Frederik Kunstner, Jacques Chen, Jonathan Wilder Lavington, and Mark Schmidt · 2023
Closest in time.