Fetching the paper…
Reading the bibliography…
We make contributions towards improving adaptive-optimizer performance.
B. T. Polyak, “Some methods of speeding up the convergence of iteration methods,” USSR Computational Mathematics and Mathematical Physics , vol. 4, pp. 1–17, 1964
1964
Earlier work this paper cites.
M. Marcus, B. Santorini, and M. A. Marcinkiewicz, “Building a large annotated corpus of english: The penn treebank,” 1993. [Online]. Available: https://catalog.ldc.upenn.edu/docs/LDC95T7/cl93.html
1993
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long Short-Term Memory,” Neural Computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
J. Duchi, E. Hazan, and Y. Singer, “Adaptive Subgradient Methods for Online Learning and Stochastic Optimization,” Journal of Machine Learning Research , vol. 12, pp. 2121–2159, 2011
2011
Earlier work this paper cites.
T. Tieleman and G. Hinton, “Lecture 6.5-RMSProp: Divide The Gradient by a Running Average of Its Recent Magnitude,” COURSERA: Neural networks for machine learning, pp. 26–31, 2012
2012
Earlier work this paper cites.
H. Sutskever, J. Martens, G. Dahl, and G. Hinton, “On the importance of initialization and momentum in deep learning,” in International conference on Machine Learning (ICML) , 2013
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep Learning,” Nature , vol. 521, pp. 436–444, 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in IEEE conference on Computer Vision and Pattern Recognition (CVPR) , 2015
2015
Earlier work this paper cites.
D. S. et al., “”Mastering the game of Go with deep neural networks and tree search,” Nature , vol. 529, no. 7587, pp. 484–489, 2016
2016
Cited alongside, same era.
T. Dozat, “Incorporating Nesterov Momentum into Adam,” in International conference on Learning Representations (ICLR) , 2016
2016
Cited alongside, same era.
2017
Cited alongside, same era.
A. C. Wilson, R. Roelofs, M. Stern, N. Srebro, and B. Recht, “The Marginal Value of Adaptive Gradient Methods in Machine Learning,” in 31st Conference on Neural Information Processing Systems (NIPS) , 2017
2017
Cited alongside, same era.
J. Zhang, S. P. Karimireddy, A. Veit, S. Kim, S. J. Reddi, S. Kumar, and S. Sra, “Why ADAM Beats SGD for Attention Models,” in submitted for review by ICLR , 2019
2019
Later among the works it cites.
I. Loshchilov and F. Hutter, “Decoupled Weight Decay Regularization,” in ICLR , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
L. Luo, Y. Xiong, Y. Liu, and X. sun, “Adaptive Gradient Methods with Dynamic Bound of Learning Rate,” in ICLR , 2019
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
2017
Cited alongside, same era.
I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville, “Improved training of wasserstein gans,” in Advances in neural information processing systems , 2017, pp. 5767–5777
2017
Cited alongside, same era.
2018
Cited alongside, same era.
M. Zaheer, S. Reddi, D. Sachan, S. Kale, and S. Kumar, “Adaptive methods for nonconvex optimization,” in Advances in neural information processing systems (NeurIPS) , 2018, p. 9793–9803
2018
Cited alongside, same era.
2020
Later among the works it cites.
2020
Later among the works it cites.
J. Zhuang, T. Tang, S. T. Y. Ding, N. Dvornek, X. Papademetris, and J. S. Duncan, “AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed Gradients,” in NeurIPS , 2020
2020
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin Transformer: Hierarchical Vision Transformer using Shifted Windows,” in International Conference on Computer Vision (ICCV) , 2021
2021
Later among the works it cites.