Fetching the paper…
Reading the bibliography…
Stochastic methods with coordinate-wise adaptive stepsize (such as RMSprop and Adam) have been widely used in training deep neural networks.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
Problem Complexity and Method Efficiency in Optimization
A. Nemirovski and D.B. Yudin · 1983
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o (1/kˆ 2)
Y. E. Nesterov · 1983
Earlier work this paper cites.
A training algorithm for optimal margin classifiers
B. E. Boser, I. M. Guyon, and V. N. Vapnik · 1992
Earlier work this paper cites.
Stability and generalization
O. Bousquet and A. Elisseeff · 2002
Earlier work this paper cites.
On the generalization ability of on-line learning algorithms
N. Cesa-Bianchi, A. Conconi, and C. Gentile · 2004
Earlier work this paper cites.
Loss functions for preference levels: Regression with discrete ordered labels
J. D. Rennie and N. Srebro · 2005
Earlier work this paper cites.
The Cambridge dictionary of statistics
B. S. Everitt · 2006
Earlier work this paper cites.
Adaptive bound optimization for online convex optimization
H. B. McMahan and M. Streeter · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Lecture 6.5 - RMSProp, COURSERA: Neural networks for machine learning, 2012
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
ADADELTA: An adaptive learning rate method
M. D. Zeiler · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
A. Graves, A. Mohamed, and G. Hinton · 2013
Cited alongside, same era.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
O. Shamir and T. Zhang · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Cited alongside, same era.
Recurrent neural network regularization
W. Zaremba, I. Sutskever, and O. Vinyals · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Normalized gradient with adaptive stepsize method for deep neural network training
A. W. Yu, Q. Lin, R. Salakhutdinov, and J. Carbonell · 2017
Later among the works it cites.
Theory of deep learning iii: Generalization properties of sgd
C. Zhang, Q. Liao, A. Rakhlin, K. Sridharan, B. Miranda, N. Golowich, and T. Poggio · 2017
Later among the works it cites.
Dissecting adam: The sign, magnitude and variance of stochastic gradients
L. Balles and P. Hennig · 2018
Later among the works it cites.
signSGD: Compressed optimisation for non-convex problems
J. Bernstein, Y. Wang, K. Azizzadenesheli, and A. Anandkumar · 2018
Later among the works it cites.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2018
Later among the works it cites.
Closing the generalization gap of adaptive gradient methods in training deep neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Layer-specific adaptive learning rates for deep networks
B. Singh, S. De, Y. Zhang, T. Goldstein, and G. Taylor · 2015
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
M. Hardt, B. Recht, and Y. Singer · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
SGDR: Stochastic gradient descent with warm restarts
I. Loshchilov and F. Hutter · 2017
Cited alongside, same era.
Pointer sentinel mixture models
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
J. Chen and Q. Gu · 2018
Later among the works it cites.
Regularizing and optimizing LSTM language models
S. Merity, N. S. Keskar, and R. Socher · 2018
Later among the works it cites.
Adagrad stepsizes: Sharp convergence over nonconvex landscapes, from any initialization
R. Ward, X. Wu, and L. Bottou · 2018
Later among the works it cites.
Adaptive methods for nonconvex optimization
M. Zaheer, S. Reddi, D. Sachan, S. Kale, and S. Kumar · 2018
Later among the works it cites.
mixup: Beyond empirical risk minimization
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz · 2018
Later among the works it cites.
Adashift: Decorrelation and convergence of adaptive learning rate methods
Z. Zhou, Q. Zhang, G. Lu, H. Wang, W. Zhang, and Y. Yu · 2018
Later among the works it cites.
On the convergence of weighted adagrad with momentum for training deep neural networks
F. Zou and L. Shen · 2018
Later among the works it cites.
A sufficient condition for convergences of adam and rmsprop
F. Zou, L. Shen, Z. Jie, W. Zhang, and W. Liu · 2018
Later among the works it cites.