Fetching the paper…
Reading the bibliography…
Adaptive gradient methods such as Adam have gained increasing popularity in deep learning optimization.
Adaptive gradient methods with dynamic bound of learning rate
Luo, L · 1902
Earlier work this paper cites.
Yang, G · 1902
Earlier work this paper cites.
On the convergence of adam and beyond
Reddi, S. J · 1904
Earlier work this paper cites.
Backward feature correction: How deep learning performs deep learning
Allen-Zhu, Z · 2001
Earlier work this paper cites.
Feature purification: How adversarial training performs robust deep learning
Allen-Zhu, Z · 2005
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J · 2011
Earlier work this paper cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Allen-Zhu, Z · 2012
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Hinton, G · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P · 2015
Earlier work this paper cites.
Analyzing tensor power method dynamics in overcomplete regime
Anandkumar, A · 2017
Earlier work this paper cites.
Improving generalization performance by switching from adam to sgd
Keskar, N. S · 2017
Cited alongside, same era.
Convolutional neural networks analyzed via convolutional sparse coding
Papyan, V · 2017
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
Wilson, A. C · 2017
Cited alongside, same era.
Dissecting adam: The sign, magnitude and variance of stochastic gradients
Balles, L · 2018
Cited alongside, same era.
signsgd: Compressed optimisation for non-convex problems
Bernstein, J · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A · 2018
Cited alongside, same era.
What can ResNet learn efficiently, going beyond kernels?
Allen-Zhu, Z · 2019
Later among the works it cites.
Beyond linearization: On quadratic and higher-order approximation of wide neural networks
Bai, Y · 2019
Later among the works it cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Cao, Y · 2019
Later among the works it cites.
Gradient descent optimizes over-parameterized deep ReLU networks
Zou, D · 2019
Later among the works it cites.
An improved analysis of training over-parameterized deep neural networks
Zou, D · 2019
Later among the works it cites.
Closing the generalization gap of adaptive gradient methods in training deep neural networks
Chen, J · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y · 2018
Cited alongside, same era.
Regularizing and optimizing lstm language models
Merity, S · 2018
Cited alongside, same era.
On the convergence of adam and beyond
Reddi, S. J · 2018
Cited alongside, same era.
Revisiting the generalization of adaptive gradient methods
Agarwal, N · 2019
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Allen-Zhu, Z
Cited in the paper.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z
Cited in the paper.
Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow relu networks
Ji, Z · 2020
Later among the works it cites.
Learning over-parametrized two-layer neural networks beyond ntk
Li, Y · 2020
Later among the works it cites.
Towards theoretically understanding why sgd generalizes better than adam in deep learning
Zhou, P · 2020
Later among the works it cites.
How much over-parameterization is sufficient to learn deep relu networks?
Chen, Z · 2021
Closest in time.