Fetching the paper…
Reading the bibliography…
Adaptive gradient methods, e.g.
Why are adaptive methods good for attention models?
Zhang, J · 1912
Earlier work this paper cites.
A stochastic approximation method
Robbins, H · 1951
Earlier work this paper cites.
Théorie de l’addition des variables aléatoires
Lévy, P · 1954
Earlier work this paper cites.
The global optimization problem. an introduction
Dixon, L. C. W · 1978
Earlier work this paper cites.
Lectures on geometric measure theory
Simon, L · 1983
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Marcus, M. P · 1993
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M · 2003
Earlier work this paper cites.
A study of gradient variance in deep learning
Faghri, F · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Ma, X · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J · 2011
Earlier work this paper cites.
First exit times of solutions of stochastic differential equations driven by multiplicative lévy noise with heavy tails
Pavlyukevich, I · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman, G. H · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I · 2013
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Bubeck, S · 2014
Earlier work this paper cites.
Report on the 11th iwslt evaluation campaign, iwslt 2014
Cettolo, M · 2014
Earlier work this paper cites.
On the properties of neural machine translation: Encoder–decoder approaches
Cho, K · 2014
Cited alongside, same era.
Generative adversarial networks
Goodfellow, I. J · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P · 2015
Cited alongside, same era.
Unsupervised representation learning with deep convolutional generative adversarial networks
Radford, A · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Russakovsky, O · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, K · 2015
Cited alongside, same era.
Averaging weights leads to wider optima and better generalization
Izmailov, P · 2018
Later among the works it cites.
Visualizing the loss landscape of neural nets
Li, H · 2018
Later among the works it cites.
Spectral normalization for generative adversarial networks
Miyato, T · 2018
Later among the works it cites.
Training tips for the transformer model
Popel, M · 2018
Later among the works it cites.
Large scale GAN training for high fidelity natural image synthesis
Brock, A · 2019
Later among the works it cites.
On the convergence of a class of adam-type algorithms for non-convex optimization
Chen, X · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Incorporating nesterov momentum into adam
Dozat, T · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Szegedy, C · 2016
Cited alongside, same era.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M · 2017
Cited alongside, same era.
Densely connected convolutional networks
Huang, G · 2017
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S · 2017
Cited alongside, same era.
He, H · 2019
Later among the works it cites.
On the variance of the adaptive learning rate and beyond
Liu, L · 2019
Later among the works it cites.
Adaptive gradient methods with dynamic bound of learning rate
Luo, L · 2019
Later among the works it cites.
A tail-index analysis of stochastic gradient noise in deep neural networks
Simsekli, U · 2019
Later among the works it cites.
Consistency regularization for generative adversarial networks
Zhang, H · 2020
Later among the works it cites.
Towards theoretically understanding why sgd generalizes better than adam in deep learning
Zhou, P · 2020
Later among the works it cites.
Adabelief optimizer: Adapting stepsizes by the belief in observed gradients
Zhuang, J · 2020
Later among the works it cites.
On stochastic moving-average estimators for non-convex optimization
Guo, Z · 2021
Closest in time.
Biadam: Fast adaptive bilevel optimization methods
Huang, F · 2021
Closest in time.
Domain-independent dominance of adaptive methods
Savarese, P · 2021
Closest in time.
Descending through a crowded valley–benchmarking deep learning optimizers
Schmidt, R. M · 2021
Closest in time.