Fetching the paper…
Reading the bibliography…
The adaptive moment estimation algorithm Adam (Kingma and Ba) is a popular optimizer in the training of deep neural networks.
”A stochastic approximation method”,
Herbert Robbins and Sutton Monro, · 1951
Earlier work this paper cites.
”Adaptive Bound Optimization for Online Convex Optimization”,
H. Brendan McMahan and Matthew Streeter, · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. (2015) · 2015
Earlier work this paper cites.
”Deep Residual Learning for Image Recognition”,
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Cited alongside, same era.
”Identity Mappings in Deep Residual Networks”,
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Cited alongside, same era.
On the convergence of Adam and beyond
Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar. (2018) · 2018
Cited alongside, same era.
”An improvement of the convergence proof of the Adam-optimizer”,
Sebastian Bock, Josef Goppold, and Martin Weiß, · 2018
Later among the works it cites.
”Closing the generalization gap of adaptive gradient methods in training deep neural networks”,
Jinghui Chen and Quanquan Gu, · 2018
Later among the works it cites.
Adaptive gradient methods with dynamic bound of learning rate
Liangchen Luo, Yuanhao Xiong, and Yan Liu. (2019) · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…