Fetching the paper…
Reading the bibliography…
Adam is widely adopted in practical applications due to its fast convergence.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Curiously fast convergence of some stochastic gradient descent algorithms. In Proceedings of the symposium on learning and data science, Paris , Vol. 8. 2624–2633
Léon Bottou. 2009 · 2009
Earlier work this paper cites.
Stochastic gradient descent tricks
Léon Bottou. 2012 · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan. 2013 · 2013
Earlier work this paper cites.
Fast convergence of stochastic gradient descent under a strong growth condition
Mark Schmidt and Nicolas Le Roux. 2013 · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Advanced Optimization: Lecture 18 Proximal methods, Monotone operators
Suvrit Sra. 2014 · 2014
Earlier work this paper cites.
Global convergence of the heavy-ball method for convex optimization. In 2015 European control conference (ECC) . IEEE, 310–315
Euhanna Ghadimi, Hamid Reza Feyzmahdavian, and Mikael Johansson. 2015 · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization. In International Conference on Learning Representations
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Unsupervised representation learning with deep convolutional generative adversarial networks
Alec Radford, Luke Metz, and Soumith Chintala. 2015 · 2015
Earlier work this paper cites.
Incorporating Nesterov momentum into Adam
Timothy Dozat. 2016 · 2016
Earlier work this paper cites.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
Saeed Ghadimi, Guanghui Lan, and Hongchao Zhang. 2016 · 2016
Earlier work this paper cites.
Attention is all you need. In Advances in neural information processing systems . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Large Scale GAN Training for High Fidelity Natural Image Synthesis. In International Conference on Learning Representations
Andrew Brock, Jeff Donahue, and Karen Simonyan. 2018 · 2018
Earlier work this paper cites.
On the convergence of a class of Adam-type algorithms for non-convex optimization
Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong. 2018 · 2018
Earlier work this paper cites.
Soham De, Anirbit Mukherjee, and Enayat Ullah. 2018 · 2018
Earlier work this paper cites.
On the Convergence of Adam and Beyond. In International Conference on Learning Representations
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar. 2018 · 2018
Cited alongside, same era.
Adaptive methods for nonconvex optimization
Manzil Zaheer, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar. 2018a · 2018
Cited alongside, same era.
Adaptive Methods for Nonconvex Optimization. In Advances in Neural Information Processing Systems , S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Eds.), Vol. 31. Curran Associates, Inc
Manzil Zaheer, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar. 2018b · 2018
Cited alongside, same era.
On the convergence of adaptive gradient methods for nonconvex optimization
Dongruo Zhou, Jinghui Chen, Yuan Cao, Yiqi Tang, Ziyan Yang, and Quanquan Gu. 2018a · 2018
Cited alongside, same era.
Adashift: Decorrelation and convergence of adaptive learning rate methods
Improved analysis of clipping algorithms for non-convex optimization
Bohang Zhang, Jikai Jin, Cong Fang, and Liwei Wang. 2020 · 2020
Later among the works it cites.
Towards Practical Adam: Non-Convexity, Convergence Theory, and Mini-Batch Acceleration
Congliang Chen, Li Shen, Fangyu Zou, and Wei Liu. 2021 · 2021
Later among the works it cites.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy M Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar. 2021 · 2021
Later among the works it cites.
A Novel Convergence Analysis for Algorithms of the Adam Family
Zhishuai Guo, Yi Xu, Wotao Yin, Rong Jin, and Tianbao Yang. 2021 · 2021
Later among the works it cites.
Super-Adam: faster and universal framework of adaptive gradients
Feihu Huang, Junyi Li, and Heng Huang. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhiming Zhou, Qingru Zhang, Guansong Lu, Hongwei Wang, Weinan Zhang, and Yong Yu. 2018b · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of NAACL-HLT . 4171–4186
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Adaptive gradient methods with dynamic bound of learning rate
Liangchen Luo, Yuanhao Xiong, Yan Liu, and Xu Sun. 2019 · 2019
Cited alongside, same era.
On the convergence of Adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar. 2019 · 2019
Cited alongside, same era.
Fast and faster convergence of SGD for over-parameterized models and an accelerated perceptron. In The 22nd International Conference on Artificial Intelligence and Statistics . PMLR, 1195–1204
Sharan Vaswani, Francis Bach, and Mark Schmidt. 2019 · 2019
Cited alongside, same era.
Why Adam beats SGD for attention models
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank J Reddi, Sanjiv Kumar, and Suvrit Sra. 2019b · 2019
Cited alongside, same era.
Fangyu Zou, Li Shen, Zequn Jie, Weizhong Zhang, and Wei Liu. 2019 · 2019
Cited alongside, same era.
A Simple Convergence Proof of Adam and AdaGrad
Alexandre Défossez, Léon Bottou, Francis Bach, and Nicolas Usunier. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
Non-convex distributionally robust optimization: Non-asymptotic analysis
Jikai Jin, Bohang Zhang, Haiyang Wang, and Liwei Wang. 2021 · 2021
Later among the works it cites.
RMSprop converges with proper hyper-parameter. In International Conference on Learning Representations
Naichen Shi, Dawei Li, Mingyi Hong, and Ruoyu Sun. 2021 · 2021
Later among the works it cites.
Worst-case complexity of cyclic coordinate descent: O (nˆ 2) O (n 2) gap with randomized version
Ruoyu Sun and Yinyu Ye. 2021 · 2021
Later among the works it cites.
Smg: A shuffling gradient-based method with momentum. In International Conference on Machine Learning . PMLR, 10379–10389
Trang H Tran, Lam M Nguyen, and Quoc Tran-Dinh. 2021 · 2021
Later among the works it cites.
Robustness to Unbounded Smoothness of Generalized SignSGD
Michael Crawshaw, Mingrui Liu, Francesco Orabona, Wei Zhang, and Zhenxun Zhuang. 2022 · 2022
Closest in time.
The Power of Adaptivity in SGD: Self-Tuning Step Sizes with Unbounded Gradients and Affine Variance
Matthew Faw, Isidoros Tziotis, Constantine Caramanis, Aryan Mokhtari, Sanjay Shakkottai, and Rachel Ward. 2022 · 2022
Closest in time.
Normalized/Clipped SGD with Perturbation for Differentially Private Non-Convex Optimization
Xiaodong Yang, Huishuai Zhang, Wei Chen, and Tie-Yan Liu. 2022 · 2022
Closest in time.
Adam Can Converge Without Any Modification on Update Rules
Yushun Zhang, Congliang Chen, Naichen Shi, Ruoyu Sun, and Zhi-Quan Luo. 2022 · 2022
Closest in time.
Convergence of Adam Under Relaxed Assumptions
Haochuan Li, Ali Jadbabaie, and Alexander Rakhlin. 2023 · 2023
Closest in time.
Why Transformers Need Adam: A Hessian Perspective
Yushun Zhang, Congliang Chen, Tian Ding, Ziniu Li, Ruoyu Sun, and Zhi-Quan Luo. 2024 · 2024
Closest in time.