Fetching the paper…
Reading the bibliography…
Very little is known about the training dynamics of adaptive gradient methods like Adam in deep learning.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Incorporating nesterov momentum into adam
Timothy Dozat · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Why momentum really works
Gabriel Goh · 2017
Earlier work this paper cites.
Improving generalization performance by switching from adam to sgd
Nitish Shirish Keskar and Richard Socher · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Earlier work this paper cites.
Dissecting adam: The sign, magnitude and variance of stochastic gradients
Lukas Balles and Philipp Hennig · 2018
Earlier work this paper cites.
Convergence and dynamical behavior of the adam algorithm for non-convex stochastic optimization
Anas Barakat and Pascal Bianchi · 2018
Earlier work this paper cites.
Closing the generalization gap of adaptive gradient methods in training deep neural networks
Jinghui Chen, Dongruo Zhou, Yiqi Tang, Ziyan Yang, Yuan Cao, and Quanquan Gu · 2018
Earlier work this paper cites.
On the convergence of a class of adam-type algorithms for non-convex optimization
Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong · 2018
Earlier work this paper cites.
On the relation between the sharpest directions of dnn loss and the sgd step length
Stanisław Jastrzębski, Zachary Kenton, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2018
Earlier work this paper cites.
Step size matters in deep learning
Kamil Nar and Shankar Sastry · 2018
Earlier work this paper cites.
Training tips for the transformer model
Martin Popel and Ondřej Bojar · 2018
Cited alongside, same era.
On the convergence of adam and beyond
Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Cited alongside, same era.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, and Weinan E · 2018
Cited alongside, same era.
On the convergence of adaptive gradient methods for nonconvex optimization
Dongruo Zhou, Jinghui Chen, Yuan Cao, Yiqi Tang, Ziyan Yang, and Quanquan Gu · 2018
Cited alongside, same era.
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Later among the works it cites.
Understanding the difficulty of training transformers
Liyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen, and Jiawei Han · 2020
Later among the works it cites.
A qualitative study of the dynamic behavior for adaptive gradient algorithms
Chao Ma, Lei Wu, et al · 2020
Later among the works it cites.
Linear convergence of adaptive stochastic gradient descent
Yuege Xie, Xiaoxia Wu, and Rachel Ward · 2020
Later among the works it cites.
Adai: Separating the effects of adaptive learning rate and momentum inertia
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Niv Giladi, Mor Shpigel Nacson, Elad Hoffer, and Daniel Soudry · 2019
Cited alongside, same era.
On the convergence of stochastic gradient descent with adaptive stepsizes
Xiaoyu Li and Francesco Orabona · 2019
Cited alongside, same era.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2019
Cited alongside, same era.
Adagrad stepsizes: Sharp convergence over nonconvex landscapes
Rachel Ward, Xiaoxia Wu, and Leon Bottou · 2019
Cited alongside, same era.
Disentangling adaptive gradient methods from learning rates
Naman Agarwal, Rohan Anil, Elad Hazan, Tomer Koren, and Cyril Zhang · 2020
Cited alongside, same era.
Google scholar reveals its most influential papers for 2020, Jul 2020
Bec Crew · 2020
Cited alongside, same era.
A general system of differential equations to model first-order adaptive algorithms
André Belotto da Silva and Maxime Gazeau · 2020
Cited alongside, same era.
Zeke Xie, Xinrui Wang, Huishuai Zhang, Issei Sato, and Masashi Sugiyama · 2020
Later among the works it cites.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu · 2020
Later among the works it cites.
Towards theoretically understanding why sgd generalizes better than adam in deep learning
Pan Zhou, Jiashi Feng, Chao Ma, Caiming Xiong, Steven Chu Hong Hoi, et al · 2020
Later among the works it cites.
When vision transformers outperform resnets without pre-training or strong data augmentations
Xiangning Chen, Cho-Jui Hsieh, and Boqing Gong · 2021
Later among the works it cites.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy M. Cohen, Simran Kaur, Yuanzhi Li, J. Zico Kolter, and Ameet S. Talwalkar · 2021
Later among the works it cites.
A loss curvature perspective on training instability in deep learning
Justin Gilmer, Behrooz Ghorbani, Ankush Garg, Sneha Kudugunta, Behnam Neyshabur, David Cardoze, George Dahl, Zachary Nado, and Orhan Firat · 2021
Later among the works it cites.
Catastrophic fisher explosion: Early phase fisher matrix impacts generalization
Stanislaw Jastrzebski, Devansh Arpit, Oliver Astrand, Giancarlo B Kerg, Huan Wang, Caiming Xiong, Richard Socher, Kyunghyun Cho, and Krzysztof J Geras · 2021
Later among the works it cites.
The implicit bias of minima stability: A view from function space
Rotem Mulayoff, Tomer Michaeli, and Daniel Soudry · 2021
Later among the works it cites.
The implicit bias for adaptive optimization algorithms on homogeneous neural networks
Bohan Wang, Qi Meng, Wei Chen, and Tie-Yan Liu · 2021
Later among the works it cites.
Understanding the unstable convergence of gradient descent
Kwangjun Ahn, Jingzhao Zhang, and Suvrit Sra · 2022
Closest in time.
Implicit jacobian regularization weighted with impurity of probability output, 2022
Sungyoon Lee, Jinseong Park, and Jaewook Lee · 2022
Closest in time.
The multiscale structure of neural network loss functions: The effect on optimization and origin
Chao Ma, Lei Wu, and Lexing Ying · 2022
Closest in time.