Fetching the paper…
Reading the bibliography…
The success of the Adam optimizer on a wide array of architectures has made it the default in settings where stochastic gradient descent (SGD) performs poorly.
“A method of solving a convex programming problem with convergence rate O ( k 2 ) O(k^{2}) ” English translation of Russian, in Doklady Adademii Nauk SSSR 269, 543–547 (1983)
Yurii. Nesterov · 1983
Earlier work this paper cites.
“Gradient-Based Learning Applied to Document Recognition”
Yann LeCun, Léon Bottou, Yoshua Bengio and Patrick Haffner · 1998
Earlier work this paper cites.
“Nonlinear Programming”
Dimitri. Bertsekas · 1999
Earlier work this paper cites.
“The Geometry of Sign Gradient Descent” arXiv/2002.08056, 2020
Lukas Balles, Fabian Pedregosa and Nicolas Roux · 2002
Earlier work this paper cites.
“Building a Large Annotated Corpus of English: The Penn Treebank”
Mitchell. Marcus, Beatrice Santorini and Mary Marcinkiewicz · 2004
Earlier work this paper cites.
“Learning Multiple Layers of Features from Tiny Images”, 2012
Alex Krizhevsky · 2012
Earlier work this paper cites.
“RMSPROP: Divide the gradient by a running average of its recent magnitude” Lecture notes
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
http://www.cs.toronto.edu/˜tijmen/csc321/slides/lecture_slides_lec6.pdf , 2012
2012
Earlier work this paper cites.
“On the difficulty of training recurrent neural networks”
Razvan Pascanu, Tomás Mikolov and Yoshua Bengio · 2013
Earlier work this paper cites.
“From Averaging to Acceleration, There is Only a Step-size”
Nicolas Flammarion and Francis. Bach · 2015
Earlier work this paper cites.
“Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift”
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization”
Diederik. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
“Optimizing Neural Networks with Kronecker-factored Approximate Curvature”
James Martens and Roger. Grosse · 2015
Earlier work this paper cites.
Jimmy Ba, Jamie Kiros and Geoffrey. Hinton · 2016
Earlier work this paper cites.
“Deep Residual Learning for Image Recognition”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2016
Earlier work this paper cites.
“SQuAD: 100,000+ Questions for Machine Comprehension of Text”
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev and Percy Liang · 2016
Earlier work this paper cites.
“Pointer Sentinel Mixture Models”
Stephen Merity, Caiming Xiong, James Bradbury and Richard Socher · 2017
Earlier work this paper cites.
“Attention is All you Need”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan. Gomez, Lukasz Kaiser and Illia Polosukhin · 2017
Cited alongside, same era.
“Dissecting Adam: The Sign, Magnitude and Variance of Stochastic Gradients”
Lukas Balles and Philipp Hennig · 2018
Cited alongside, same era.
“SIGNSGD: Compressed Optimisation for Non-Convex Problems”
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli and Animashree Anandkumar · 2018
Cited alongside, same era.
“Lectures on Convex Optimization” 87
Yurii. Nesterov · 2018
Cited alongside, same era.
“On the Convergence of Adam and Beyond”
Sashank. Reddi, Satyen Kale and Sanjiv Kumar · 2018
Cited alongside, same era.
“Memory Efficient Adaptive Optimization”
Rohan Anil, Vineet Gupta, Tomer Koren and Yoram Singer · 2019
Cited alongside, same era.
“On the Generalization Benefit of Noise in Stochastic Gradient Descent”
Samuel. Smith, Erich Elsen and Soham De · 2020
Later among the works it cites.
“Transformers: State-of-the-Art Natural Language Processing”
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest and Alexander. Rush · 2020
Later among the works it cites.
“Large Batch Optimization for Deep Learning: Training BERT in 76 minutes”
Yang You, Jing Li, Sashank. Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer and Cho-Jui Hsieh · 2020
Later among the works it cites.
“Improved Analysis of Clipping Algorithms for Non-convex Optimization”
Bohang Zhang, Jikai Jin, Cong Fang and Liwei Wang · 2020
Later among the works it cites.
“Why Gradient Clipping Accelerates Training: A Theoretical Justification for Adaptivity”
Jingzhao Zhang, Tianxing He, Suvrit Sra and Ali Jadbabaie · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Transformer-XL: Attentive Language Models beyond a Fixed-Length Context”
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime. Carbonell, Quoc Le and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
“Convergence of Gradient Descent on Separable Data”
Mor Nacson, Jason. Lee, Suriya Gunasekar, Pedro Savarese, Nathan Srebro and Daniel Soudry · 2019
Cited alongside, same era.
“PyTorch: An Imperative Style, High-Performance Deep Learning Library”
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai and Soumith Chintala · 2019
Cited alongside, same era.
“DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter”
Victor Sanh, Lysandre Debut, Julien Chaumond and Thomas Wolf · 2019
Cited alongside, same era.
“Measuring the Effects of Data Parallelism on Neural Network Training”
Christopher. Shallue, Jaehoon Lee, Joseph. Antognini, Jascha Sohl-Dickstein, Roy Frostig and George. Dahl · 2019
Cited alongside, same era.
“Which Algorithmic Choices Matter at Which Batch Sizes? Insights From a Noisy Quadratic Model”
Guodong Zhang, Lala Li, Zachary Nado, James Martens, Sushant Sachdeva, George. Dahl, Christopher. Shallue and Roger. Grosse · 2019
Cited alongside, same era.
“Why are Adaptive Methods Good for Attention Models?”
Jingzhao Zhang, Sai Karimireddy, Andreas Veit, Seungyeon Kim, Sashank. Reddi, Sanjiv Kumar and Suvrit Sra · 2020
Later among the works it cites.
“Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability”
Jeremy. Cohen, Simran Kaur, Yuanzhi Li, J. Kolter and Ameet Talwalkar · 2021
Later among the works it cites.
Zachary Nado, Justin Gilmer, Christopher. Shallue, Rohan Anil and George. Dahl · 2021
Later among the works it cites.
“Understanding Gradient Clipping In Incremental Gradient Methods”
Jiang Qian, Yuren Wu, Bojin Zhuang, Shaojun Wang and Jing Xiao · 2021
Later among the works it cites.
“Stochastic Sign Descent Methods: New Algorithms and Better Theory”
Mher Safaryan and Peter Richtárik · 2021
Later among the works it cites.
“Efficient Estimators for Heavy-Tailed Machine Learning” OpenReview:5K8ZG9twKY, 2021
Vishwak Srinivasan, Adarsh Prasad, Sivaraman Balakrishnan and Pradeep Ravikumar · 2021
Later among the works it cites.
“Robustness to Unbounded Smoothness of Generalized SignSGD”
Michael Crawshaw, Mingrui Liu, Francesco Orabona, Wei Zhang and Zhenxun Zhuang · 2022
Later among the works it cites.
“A Simple Convergence Proof of Adam and Adagrad”
Alexandre Défossez, Leon Bottou, Francis Bach and Nicolas Usunier · 2022
Later among the works it cites.
“Stochastic Training is Not Necessary for Generalization”
Jonas Geiping, Micah Goldblum, Phil Pope, Michael Moeller and Tom Goldstein · 2022
Later among the works it cites.
“Introduction to Online Convex Optimization”
Elad Hazan · 2022
Later among the works it cites.
“Symbolic Discovery of Optimization Algorithms” arXiv/2302.06675
Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Yao Liu, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, Yifeng Lu and Quoc. Le · 2023
Closest in time.