Fetching the paper…
Reading the bibliography…
Deep learning methods achieve state-of-the-art performance in many application scenarios.
A method for unconstrained convex minimization problem with the rate of convergence O ( 1 / k 2 ) (1/k^{2})
Y. Nesterov · 1983
Earlier work this paper cites.
Untersuchungen zu dynamischen neuronalen Netzen
S. Hochreiter · 1991
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
M. P. Marcus, M. A. Marcinkiewicz, and B. Santorini · 1993
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
S.-I. Amari · 1998
Earlier work this paper cites.
Numerical optimization
S. Wright and J. Nocedal · 1999
Earlier work this paper cites.
Solving large scale linear prediction problems using stochastic gradient descent algorithms
T. Zhang · 2004
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Primal-dual subgradient methods for convex problems
Y. Nesterov · 2009
Earlier work this paper cites.
Less regret via online conditioning
M. Streeter and H. B. McMahan · 2010
Earlier work this paper cites.
Convex Analysis and Monotone Operator Theory in Hilbert Spaces
H. H. Bauschke and P. L. Combettes · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Practical recommendations for gradient-based training of deep architectures
Y. Bengio · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
No-regret algorithms for unconstrained online convex optimization
M. Streeter and H. B. McMahan · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
ADADELTA: an adaptive learning rate method
M. D. Zeiler · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Unconstrained online linear learning in Hilbert spaces: Minimax algorithms and normal approximations
H. B. McMahan and F. Orabona · 2014
Later among the works it cites.
Simultaneous model selection and optimization through parameter-free stochastic learning
F. Orabona · 2014
Later among the works it cites.
Recurrent neural network regularization
W. Zaremba, I. Sutskever, and O. Vinyals · 2014
Later among the works it cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng · 2015
Later among the works it cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Johnson and T. Zhang · 2013
Cited alongside, same era.
Dimension-free exponentiated gradient
F. Orabona · 2013
Cited alongside, same era.
Normalized online learning
S. Ross, P. Mineiro, and J. Langford · 2013
Cited alongside, same era.
No more pesky learning rates
T. Schaul, S. Zhang, and Y. LeCun · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Cited alongside, same era.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Cited alongside, same era.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner
Cited in the paper.
Later among the works it cites.
Scale-free algorithms for online linear optimization
F. Orabona and D. Pal · 2015
Later among the works it cites.
A generalized online mirror descent with applications to classification and regression
F. Orabona, K. Crammer, and N. Cesa-Bianchi · 2015
Later among the works it cites.
Online convex optimization with unconstrained domains and losses
A. Cutkosky and K. A. Boahen · 2016
Later among the works it cites.
Gradient descent learns linear dynamical systems
M. Hardt, T. Ma, and B. Recht · 2016
Later among the works it cites.
Coin betting and parameter-free online learning
F. Orabona and D. Pal · 2016
Later among the works it cites.
Online learning without prior information
A. Cutkosky and K. Boahen · 2017
Closest in time.