Fetching the paper…
Reading the bibliography…
Adam is one of the most popular optimization algorithms in deep learning.
Problem complexity and method efficiency in optimization (as nemirovsky and db yudin)
Charles Blair · 1985
Earlier work this paper cites.
Complexity issues in global optimization: a survey
Stephen A Vavasis · 1995
Earlier work this paper cites.
Gradient convergence in gradient methods with errors
Dimitri P Bertsekas and John N Tsitsiklis · 2000
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Lecture 6e rmsprop: Divide the gradient by a running average of its recent magnitude, 2012
Geoffrey Hinton, Nitish Srivastava, and Kevin Swersky · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Diederik P. Kingma and Max Welling · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
On the convergence of adam and beyond
Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Earlier work this paper cites.
On the convergence of adaptive gradient methods for nonconvex optimization
Dongruo Zhou, Jinghui Chen, Yuan Cao, Yiqi Tang, Ziyan Yang, and Quanquan Gu · 2018
Earlier work this paper cites.
Weighted adagrad with unified momentum
Fangyu Zou, Li Shen, Zequn Jie, Ju Sun, and Wei Liu · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
A sufficient condition for convergences of adam and rmsprop
Fangyu Zou, Li Shen, Zequn Jie, Weizhong Zhang, and Wei Liu · 2019
Cited alongside, same era.
On the convergence of a class of adam-type algorithms for non-convex optimization
Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong · 2019
Cited alongside, same era.
Adashift: Decorrelation and convergence of adaptive learning rate methods
Zhiming Zhou, Qingru Zhang, Guansong Lu, Hongwei Wang, Weinan Zhang, and Yong Yu · 2019
Cited alongside, same era.
The complexity of finding stationary points with stochastic gradient descent
Yoel Drori and Ohad Shamir · 2020
Later among the works it cites.
Nvae: A deep hierarchical variational autoencoder
Arash Vahdat and Jan Kautz · 2020
Later among the works it cites.
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
On the convergence of step decay step-size for stochastic optimization
Xiaoyu Wang, Sindri Magnússon, and Mikael Johansson · 2021
Later among the works it cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the convergence of stochastic gradient descent with adaptive stepsizes
Xiaoyu Li and Francesco Orabona · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Openwebtext corpus
Aaron Gokaslan and Vanya Cohen · 2019
Cited alongside, same era.
Random shuffling beats sgd after finite epochs
Jeff Haochen and Suvrit Sra · 2019
Cited alongside, same era.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
Rmsprop converges with proper hyper-parameter
Naichen Shi, Dawei Li, Mingyi Hong, and Ruoyu Sun · 2020
Cited alongside, same era.
Later among the works it cites.
Stable-baselines3: Reliable reinforcement learning implementations
Antonin Raffin, Ashley Hill, Adam Gleave, Anssi Kanervisto, Maximilian Ernestus, and Noah Dormann · 2021
Later among the works it cites.
Adam can converge without any modification on update rules
Yushun Zhang, Congliang Chen, Naichen Shi, Ruoyu Sun, and Zhi-Quan Luo · 2022
Later among the works it cites.
Bohan Wang, Yushun Zhang, Huishuai Zhang, Qi Meng, Zhi-Ming Ma, Tie-Yan Liu, and Wei Chen · 2022
Later among the works it cites.
A simple convergence proof of adam and adagrad
Alexandre Défossez, Leon Bottou, Francis Bach, and Nicolas Usunier · 2022
Later among the works it cites.
Better theory for SGD in the nonconvex world
Ahmed Khaled and Peter Richtárik · 2023
Later among the works it cites.
Convergence of adam under relaxed assumptions
Haochuan Li, Ali Jadbabaie, and Alexander Rakhlin · 2023
Later among the works it cites.
Closing the gap between the upper bound and lower bound of adam’s iteration complexity
Bohan Wang, Jingwen Fu, Huishuai Zhang, Nanning Zheng, and Wei Chen · 2023
Later among the works it cites.