Fetching the paper…
Reading the bibliography…
It is known that the learning rate is the most important hyper-parameter to tune for training deep neural networks.
A method of solving a convex programming problem with convergence rate o (1/k2)
Y. Nesterov · 1983
Earlier work this paper cites.
Adaptive stepsizes for recursive estimation with applications in approximate dynamic programming
A. P. George and W. B. Powell · 2006
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Neural Networks: Tricks of the Trade
Y. Bengio · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
T. Schaul, S. Zhang, and Y. LeCun · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
M. D. Zeiler · 2012
Earlier work this paper cites.
Hot swapping for online adaptation of optimization hyperparameters
K. Bache, D. DeCoste, and P. Smyth · 2014
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Earlier work this paper cites.
Towards end-to-end speech recognition with recurrent neural networks
A. Graves and N. Jaitly · 2014
Cited alongside, same era.
Adasecant: Robust adaptive secant method for stochastic gradient
C. Gulcehre and Y. Bengio · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Cited alongside, same era.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2014
Cited alongside, same era.
An empirical evaluation of deep learning on highway driving
B. Huval, T. Wang, S. Tandon, J. Kiske, W. Song, J. Pazhayampallil, M. Andriluka, R. Cheng-Yue, F. Mujica, A. Coates, et al · 2015
Closest in time.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Closest in time.
Adam: a method for stochastic optimization
D. Kingma and J. Lei-Ba · 2015
Closest in time.
ImageNet Large Scale Visual Recognition Challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Closest in time.
Identity mappings in deep residual networks
K. He, X. Zhang, S. Ren, and J. Sun · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deepface: Closing the gap to human-level performance in face verification
Y. Taigman, M. Yang, M. Ranzato, and L. Wolf · 2014
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2014
Cited alongside, same era.
The effects of hyperparameters on sgd training of neural networks
T. M. Breuel · 2015
Cited alongside, same era.
Rmsprop and equilibrated adaptive learning rates for non-convex optimization
Y. N. Dauphin, H. de Vries, J. Chung, and Y. Bengio · 2015
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Closest in time.
Densely connected convolutional networks
G. Huang, Z. Liu, and K. Q. Weinberger · 2016
Closest in time.
Deep networks with stochastic depth
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Weinberger · 2016
Closest in time.
Sgdr: Stochastic gradient descent with restarts
I. Loshchilov and F. Hutter · 2016
Closest in time.
An overview of gradient descent optimization algorithms
S. Ruder · 2016
Closest in time.