Fetching the paper…
Reading the bibliography…
Adam has been widely adopted for training deep neural networks due to less hyperparameter tuning and remarkable performance.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Proximité et dualité dans un espace hilbertien
J.-J. Moreau · 1965
Earlier work this paper cites.
Régularisation d’inéquations variationnelles par approximations successives. rev. française informat
B. Martinet · 1970
Earlier work this paper cites.
Monotone operators and the proximal point algorithm
R. T. Rockafellar · 1976
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
A. S. Nemirovsky and D. Yudin · 1983
Earlier work this paper cites.
A simple weight decay can improve generalization
A. Krogh and J. Hertz · 1992
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi · 1994
Earlier work this paper cites.
Using weight decay to optimize the generalization ability of a perceptron
S. Bos and E. Chug · 1996
Earlier work this paper cites.
Continuous and discrete-time nonlinear gradient descent: Relative loss bounds and convergence
M. K. Warmuth and A. K. Jagota · 1997
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
A. Beck and M. Teboulle · 2003
Earlier work this paper cites.
Convex Optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87
Y. Nesterov · 2004
Earlier work this paper cites.
Numerical Optimization
J. Nocedal and S. J. Wright · 2006
Earlier work this paper cites.
Improved second-order bounds for prediction with expert advice
N. Cesa-Bianchi, Y. Mansour, and G. Stoltz · 2007
Earlier work this paper cites.
Efficient online and batch learning using forward backward splitting
J. Duchi and Y. Singer · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Adaptive bound optimization for online convex optimization
H. B. McMahan and M. J. Streeter · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. C. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
ADADELTA: an adaptive learning rate method
M. D. Zeiler · 2012
Cited alongside, same era.
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2013
Cited alongside, same era.
Simultaneous model selection and optimization through parameter-free stochastic learning
F. Orabona · 2014
Cited alongside, same era.
Proximal algorithms
N. Parikh and S. Boyd · 2014
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Preconditioned stochastic gradient descent
X.-L. Li · 2018
Later among the works it cites.
Adaptive gradient methods with dynamic bound of learning rate
L. Luo, Y. Xiong, Y. Liu, and X. Sun · 2018
Later among the works it cites.
On the convergence of Adam and beyond
S. J. Reddi, S. Kale, and S. Kumar · 2018
Later among the works it cites.
On the convergence of adaptive gradient methods for nonconvex optimization, 2018
D. Zhou, Y. Tang, Z. Yang, Y. Cao, and Q. Gu · 2018
Later among the works it cites.
Stochastic (approximate) proximal point methods: Convergence, optimality, and adaptivity
H. Asi and J. C. Duchi · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Cited alongside, same era.
Deeply-supervised nets
C. Lee, S. Xie, P. Gallagher, Z. Zhang, and Z. Tu · 2015
Cited alongside, same era.
Scale-free online learning
F. Orabona and D. Pál · 2015
Cited alongside, same era.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
I. Gitman and B. Ginsburg · 2017
Cited alongside, same era.
Later among the works it cites.
Scaling object detection by transferring classification weights
J. Kuen, F. Perazzi, Z. Lin, J. Zhang, and Y.-P. Tan · 2019
Later among the works it cites.
Dense classification and implanting for few-shot learning
Y. Lifchitz, Y. Avrithis, S. Picard, and A. Bursuc · 2019
Later among the works it cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Later among the works it cites.
EvalNorm: Estimating batch normalization statistics for evaluation
S. Singh and A. Shrivastava · 2019
Later among the works it cites.
EfficientNet: Rethinking model scaling for convolutional neural networks
M. Tan and Quoc V. Le · 2019
Later among the works it cites.
Fixup initialization: Residual learning without normalization
H. Zhang, Y. N. Dauphin, and T. Ma · 2019
Later among the works it cites.
Disentangling adaptive gradient methods from learning rates
N. Agarwal, R. Anil, E. Hazan, T. Koren, and C. Zhang · 2020
Later among the works it cites.
Understanding decoupled and early weight decay
J. Bjorck, K. Q. Weinberger, and C. P. Gomes · 2020
Later among the works it cites.
End-to-end object detection with transformers
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko · 2020
Later among the works it cites.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2020
Later among the works it cites.
Batch normalization biases residual blocks towards the identity function in deep networks
S. De and S. Smith · 2020
Later among the works it cites.
High-performance large-scale image recognition without normalization
A. Brock, S. De, S. L. Smith, and K. Simonyan · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Later among the works it cites.