Fetching the paper…
Reading the bibliography…
Adam is one of the most influential adaptive stochastic algorithms for training deep neural networks, which has been pointed out to be divergent even in the simple convex setting via a few simple counterexamples.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1985
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A Krizhevsky · 2009
Earlier work this paper cites.
Mnist handwritten digit database. 2010
Yann LeCun, Corinna Cortes, and Christopher JC Burges · 2010
Earlier work this paper cites.
Distributed delayed stochastic optimization
Alekh Agarwal and John C Duchi · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Geoffrey Hinton, Nitish Srivastava, and Kevin Swersky · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Alex Krizhevsky · 2014
Earlier work this paper cites.
Efficient mini-batch training for stochastic optimization
Mu Li, Tong Zhang, Yuqiang Chen, and Alexander J Smola · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Incorporating Nesterov momentum into Adam
Timothy Dozat · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Non-convex optimization for machine learning
Prateek Jain, Purushottam Kar, et al · 2017
Earlier work this paper cites.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Xiangru Lian, Ce Zhang, Huan Zhang, Cho-Jui Hsieh, Wei Zhang, and Ji Liu · 2017
Earlier work this paper cites.
Variants of RMSProp and Adagrad with logarithmic regret bounds
Mahesh Chandra Mukkamala and Matthias Hein · 2017
Cited alongside, same era.
Large batch training of convolutional networks
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Cited alongside, same era.
Dissecting Adam: The sign, magnitude and variance of stochastic gradients
Lukas Balles and Philipp Hennig · 2018
Cited alongside, same era.
Amitabh Basu, Soham De, Anirbit Mukherjee, and Enayat Ullah · 2018
Cited alongside, same era.
signSGD: Compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar · 2018
Cited alongside, same era.
Adaptive gradient methods with dynamic bound of learning rate
Liangchen Luo, Yuanhao Xiong, Yan Liu, and Xu Sun · 2019
Later among the works it cites.
Dadam: A consensus-based distributed adaptive gradient method for online optimization
Parvin Nazari, Davoud Ataee Tarzanagh, and George Michailidis · 2019
Later among the works it cites.
Local adaalter: Communication-efficient stochastic gradient descent with adaptive learning rates
Cong Xie, Oluwasanmi Koyejo, Indranil Gupta, and Haibin Lin · 2019
Later among the works it cites.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Cited alongside, same era.
Nostalgic Adam: Weighing more of the past gradients when designing the adaptive learning rate
Haiwen Huang, Chang Wang, and Bin Dong · 2018
Cited alongside, same era.
On the convergence of Adam and beyond
Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
Sparsified sgd with memory
Sebastian U Stich, Jean-Baptiste Cordonnier, and Martin Jaggi · 2018
Cited alongside, same era.
AdaGrad stepsizes: Sharp convergence over nonconvex landscapes, from any initialization
Rachel Ward, Xiaoxia Wu, and Leon Bottou · 2018
Cited alongside, same era.
A unified analysis of stochastic momentum methods for deep learning
Yan Yan, Tianbao Yang, Zhe Li, Qihang Lin, and Yi Yang · 2018
Cited alongside, same era.
Adaptive methods for nonconvex optimization
Manzil Zaheer, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
Hao Yu, Rong Jin, and Sen Yang · 2019
Later among the works it cites.
Why are adaptive methods good for attention models?
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank J Reddi, Sanjiv Kumar, and Suvrit Sra · 2019
Later among the works it cites.
A sufficient condition for convergences of adam and rmsprop
Fangyu Zou, Li Shen, Zequn Jie, Weizhong Zhang, and Wei Liu · 2019
Later among the works it cites.
Convergence rates of a momentum algorithm with bounded adaptive step size for nonconvex optimization
Anas Barakat and Pascal Bianchi · 2020
Later among the works it cites.
Distributed online optimization via gradient tracking with adaptive momentum
Guido Carnevale, Francesco Farina, Ivano Notarnicola, and Giuseppe Notarstefano · 2020
Later among the works it cites.
Toward communication efficient adaptive gradient method
Xiangyi Chen, Xiaoyun Li, and Ping Li · 2020
Later among the works it cites.
A simple convergence proof of adam and adagrad
Alexandre Défossez, Léon Bottou, Francis Bach, and Nicolas Usunier · 2020
Later among the works it cites.
Adaptive federated optimization
Sashank Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett, Keith Rush, Jakub Konečnỳ, Sanjiv Kumar, and H Brendan McMahan · 2020
Later among the works it cites.
Large batch optimization for object detection: Training coco in 12 minutes
Tong Wang, Yousong Zhu, Chaoyang Zhao, Wei Zeng, Yaowei Wang, Jinqiao Wang, and Ming Tang · 2020
Later among the works it cites.
Quantized adam with error feedback
Congliang Chen, Li Shen, Haozhi Huang, and Wei Liu · 2021
Closest in time.
Efficient-adam: Communication-efficient distributed adam with complexity analysis
Congliang Chen, Li Shen, Wei Liu, and Zhi-Quan Luo · 2022
Closest in time.