Fetching the paper…
Reading the bibliography…
While stochastic gradient descent (SGD) is still the most popular optimization algorithm in deep learning, adaptive algorithms such as Adam have established empirical advantages over SGD in some deep learning applications such as training transformers.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Minimization of functions having Lipschitz continuous first partial derivatives
Larry Armijo · 1966
Earlier work this paper cites.
Introduction to optimization. 1987
Boris T Polyak · 1987
Earlier work this paper cites.
Stochastic gradient learning in neural networks
Léon Bottou et al · 1991
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Parallel data, tools and interfaces in OPUS
Jörg Tiedemann · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman, Geoffrey Hinton, et al · 2012
Earlier work this paper cites.
Estimation, optimization, and parallelism when data is sparse
John Duchi, Michael I Jordan, and Brendan McMahan · 2013
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
Iteration complexity of randomized block-coordinate descent methods for minimizing a composite function
Peter Richtárik and Martin Takáč · 2014
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Sébastien Bubeck · 2015
Earlier work this paper cites.
Beyond convexity: Stochastic quasi-convex optimization
Elad Hazan, Kfir Levy, and Shai Shalev-Shwartz · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Coordinate descent algorithms
Stephen J Wright · 2015
Earlier work this paper cites.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
The power of normalization: Faster evasion of saddle points
Kfir Y Levy · 2016
Cited alongside, same era.
Second-order optimization for neural networks
James Martens · 2016
Cited alongside, same era.
A primer on coordinate descent algorithms
Hao-Jun Michael Shi, Shenyinying Tu, Yangyang Xu, and Wotao Yin · 2016
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
signSGD: Compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar · 2018
Cited alongside, same era.
Optimization methods for large-scale machine learning
AdaGrad stepsizes: Sharp convergence over nonconvex landscapes
Rachel Ward, Xiaoxia Wu, and Leon Bottou · 2019
Later among the works it cites.
A simple convergence proof of Adam and Adagrad
Alexandre Défossez, Léon Bottou, Francis Bach, and Nicolas Usunier · 2020
Later among the works it cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Later among the works it cites.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
Jingzhao Zhang, Tianxing He, Suvrit Sra, and Ali Jadbabaie · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Cited alongside, same era.
Soham De, Anirbit Mukherjee, and Enayat Ullah · 2018
Cited alongside, same era.
Accelerating greedy coordinate descent methods
Haihao Lu, Robert Freund, and Vahab Mirrokni · 2018
Cited alongside, same era.
On the convergence of Adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Cited alongside, same era.
On the convergence of adaptive gradient methods for nonconvex optimization
Dongruo Zhou, Jinghui Chen, Yuan Cao, Yiqi Tang, Ziyan Yang, and Quanquan Gu · 2018
Cited alongside, same era.
On the convergence of weighted adagrad with momentum for training deep neural networks
Fangyu Zou and Li Shen · 2018
Cited alongside, same era.
Later among the works it cites.
Why are adaptive methods good for attention models?
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank Reddi, Sanjiv Kumar, and Suvrit Sra · 2020
Later among the works it cites.
GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, mar 2021
Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman · 2021
Later among the works it cites.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar · 2021
Later among the works it cites.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2021
Later among the works it cites.
Adaptive gradient methods at the edge of stability
Jeremy M Cohen, Behrooz Ghorbani, Shankar Krishnan, Naman Agarwal, Sourabh Medapati, Michal Badura, Daniel Suo, David Cardoze, Zachary Nado, George E Dahl, et al · 2022
Later among the works it cites.
Recent theoretical advances in non-convex optimization
Marina Danilova, Pavel Dvurechensky, Alexander Gasnikov, Eduard Gorbunov, Sergey Guminov, Dmitry Kamzolov, and Innokentiy Shibaev · 2022
Later among the works it cites.
How does adaptive optimization impact local neural network geometry?
Kaiqi Jiang, Dhruv Malik, and Yuanzhi Li · 2022
Later among the works it cites.
The stack: 3 tb of permissively licensed source code
Denis Kocetkov, Raymond Li, Loubna Ben Allal, Jia Li, Chenghao Mou, Carlos Muñoz Ferrandis, Yacine Jernite, Margaret Mitchell, Sean Hughes, Thomas Wolf, et al · 2022
Later among the works it cites.
Differentially private coordinate descent for composite empirical risk minimization
Paul Mangold, Aurélien Bellet, Joseph Salmon, and Marc Tommasi · 2022
Later among the works it cites.
Stochastic first-order learning for large-scale flexibly tied gaussian mixture model
Mohammad Pasande, Reshad Hosseini, and Babak Nadjar Araabi · 2022
Later among the works it cites.
Symbolic discovery of optimization algorithms
Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Yao Liu, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, et al · 2023
Closest in time.
High-dimensional private empirical risk minimization by greedy coordinate descent
Paul Mangold, Aurélien Bellet, Joseph Salmon, and Marc Tommasi · 2023
Closest in time.