Fetching the paper…
Reading the bibliography…
Following the introduction of Adam, several novel adaptive optimizers for deep learning have been proposed.
Parameter Adaptation in Stochastic Optimization , page 111–134
Luís B. Almeida, Thibault Langlois, Jose D. Amaral, and Alexander Plakhov · 1999
Earlier work this paper cites.
The unreasonable effectiveness of recurrent neural networks
Andrej Karpathy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
Dougal Maclaurin, David Duvenaud, and Ryan Adams · 2015
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando De Freitas · 2016
Earlier work this paper cites.
Incorporating nesterov momentum into adam
Timothy Dozat · 2016
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Forward and reverse gradient-based hyperparameter optimization
Luca Franceschi, Michele Donini, Paolo Frasconi, and Massimiliano Pontil · 2017
Earlier work this paper cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Earlier work this paper cites.
Learned optimizers that scale and generalize
Olga Wichrowska, Niru Maheswaranathan, Matthew W Hoffman, Sergio Gomez Colmenarejo, Misha Denil, Nando Freitas, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Large batch training of convolutional networks, 2017
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Earlier work this paper cites.
Online learning rate adaptation with hypergradient descent
Atilim Gunes Baydin, Robert Cornish, David Martinez Rubio, Mark Schmidt, and Frank Wood · 2018
Earlier work this paper cites.
On the convergence of adam and beyond
Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Earlier work this paper cites.
Adaptive methods for nonconvex optimization
Manzil Zaheer, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
Openwebtext corpus
Aaron Gokaslan and Vanya Cohen · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Adaptive gradient methods with dynamic bound of learning rate
Liangchen Luo, Yuanhao Xiong, and Yan Liu · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Closing the generalization gap of adaptive gradient methods in training deep neural networks, 2020
Jinghui Chen, Dongruo Zhou, Yiqi Tang, Ziyan Yang, Yuan Cao, and Quanquan Gu · 2020
Cited alongside, same era.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2021
Later among the works it cites.
Adamp: Slowing down the slowdown for momentum optimizers on scale-invariant weights, 2021
Byeongho Heo, Sanghyuk Chun, Seong Joon Oh, Dongyoon Han, Sangdoo Yun, Gyuwan Kim, Youngjung Uh, and Jung-Woo Ha · 2021
Later among the works it cites.
Meta-learning in neural networks: A survey
Timothy Hospedales, Antreas Antoniou, Paul Micaelli, and Amos Storkey · 2021
Later among the works it cites.
Descending through a crowded valley - benchmarking deep learning optimizers
Robin M Schmidt, Frank Schneider, and Philipp Hennig · 2021
Later among the works it cites.
Gradient descent: The ultimate optimizer
Kartik Chandra, Audrey Xie, Jonathan Ragan-Kelley, and Erik Meijer · 2022
Later among the works it cites.
A simple convergence proof of adam and adagrad
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the variance of the adaptive learning rate and beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han · 2020
Cited alongside, same era.
Luke Metz, Niru Maheswaranathan, C Daniel Freeman, Ben Poole, and Jascha Sohl-Dickstein · 2020
Cited alongside, same era.
Zero: Memory optimizations toward training trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Cited alongside, same era.
AutoML-zero: Evolving machine learning algorithms from scratch
Esteban Real, Chen Liang, David So, and Quoc Le · 2020
Cited alongside, same era.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2020
Cited alongside, same era.
Adabelief optimizer: Adapting stepsizes by the belief in observed gradients
Juntang Zhuang, Tommy Tang, Yifan Ding, Sekhar C Tatikonda, Nicha Dvornek, Xenophon Papademetris, and James Duncan · 2020
Cited alongside, same era.
Alexandre Défossez, Leon Bottou, Francis Bach, and Nicolas Usunier · 2022
Later among the works it cites.
Velo: Training versatile learned optimizers by scaling up
Luke Metz, James Harrison, C Daniel Freeman, Amil Merchant, Lucas Beyer, James Bradbury, Naman Agrawal, Ben Poole, Igor Mordatch, Adam Roberts, et al · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Later among the works it cites.
Symbolic discovery of optimization algorithms, 2023
Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Yao Liu, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, Yifeng Lu, and Quoc V. Le · 2023
Later among the works it cites.
Sophia: A scalable stochastic second-order optimizer for language model pre-training, 2023
Hong Liu, Zhiyuan Li, David Hall, Percy Liang, and Tengyu Ma · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models, 2023
Xingyu Xie, Pan Zhou, Huan Li, Zhouchen Lin, and Shuicheng Yan · 2023
Later among the works it cites.