Fetching the paper…
Reading the bibliography…
The move from hand-designed features to learned features in machine learning has been wildly successful.
A method of solving a convex programming problem with convergence rate o (1/k2)
Y. Nesterov · 1983
Earlier work this paper cites.
Evolutionary principles in self-referential learning; On learning how to learn: The meta-meta-… hook
J. Schmidhuber · 1987
Earlier work this paper cites.
Integer and combinatorial optimization
G. L. Nemhauser and L. A. Wolsey · 1988
Earlier work this paper cites.
Learning a synaptic learning rule
Y. Bengio, S. Bengio, and J. Cloutier · 1990
Earlier work this paper cites.
Fixed-weight networks can learn
N. E. Cotter and P. R. Conwell · 1990
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to dynamic recurrent networks
J. Schmidhuber · 1992
Earlier work this paper cites.
Adapting bias by gradient descent: An incremental version of delta-bar-delta
R. S. Sutton · 1992
Earlier work this paper cites.
A direct adaptive method for faster backpropagation learning: The RPROP algorithm
M. Riedmiller and H. Braun · 1993
Earlier work this paper cites.
A neural network that embeds its own meta-levels
J. Schmidhuber · 1993
Earlier work this paper cites.
On the search for new learning rules for ANNs
S. Bengio, Y. Bengio, and J. Cloutier · 1995
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Shifting inductive bias with success-story algorithm, adaptive levin search, and incremental self-improvement
J. Schmidhuber, J. Zhao, and M. Wiering · 1997
Earlier work this paper cites.
No free lunch theorems for optimization
D. H. Wolpert and W. G. Macready · 1997
Earlier work this paper cites.
A signal processing framework based on dynamic neural networks with application to problems in adaptation, filtering, and classification
L. A. Feldkamp and G. V. Puskorius · 1998
Cited alongside, same era.
Learning to learn
S. Thrun and L. Pratt · 1998
Cited alongside, same era.
An incremental gradient (-projection) method with momentum term and adaptive stepsize rule
P. Tseng · 1998
Cited alongside, same era.
Local gain adaptation in stochastic gradient descent
N. N. Schraudolph · 1999
Cited alongside, same era.
Fixed-weight on-line learning
A. S. Younger, P. R. Conwell, and N. E. Cotter · 1999
Cited alongside, same era.
Evolution and design of distributed learning rules
T. P. Runarsson and M. T. Jonsson · 2000
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Later among the works it cites.
neuron, 2011
T. Maley · 2011
Later among the works it cites.
Optimization with sparsity-inducing penalties
F. Bach, R. Jenatton, J. Mairal, and G. Obozinski · 2012
Later among the works it cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Later among the works it cites.
Advances in optimizing recurrent networks
Y. Bengio, N. Boulanger-Lewandowski, and R. Pascanu · 2013
Later among the works it cites.
Neural Turing machines
A. Graves, G. Wayne, and I. Danihkela · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to learn using gradient descent
S. Hochreiter, A. S. Younger, and P. R. Conwell · 2001
Cited alongside, same era.
Meta-learning with backpropagation
A. S. Younger, S. Hochreiter, and P. R. Conwell · 2001
Cited alongside, same era.
Compressed sensing
D. L. Donoho · 2006
Cited alongside, same era.
Numerical optimization
J. Nocedal and S. Wright · 2006
Cited alongside, same era.
brain-neurons, 2009
F. Bobolas · 2009
Cited alongside, same era.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Cited alongside, same era.
L. A. Gatys, A. S. Ecker, and M. Bethge · 2015
Later among the works it cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Later among the works it cites.
Optimizing neural networks with Kronecker-factored approximate curvature
J. Martens and R. Grosse · 2015
Later among the works it cites.
Learning step size controllers for robust neural network training
C. Daniel, J. Taylor, and S. Nowozin · 2016
Closest in time.
Building machines that learn and think like people
B. M. Lake, T. D. Ullman, J. B. Tenenbaum, and S. J. Gershman · 2016
Closest in time.
Meta-learning with memory-augmented neural networks
A. Santoro, S. Bartunov, M. Botvinick, D. Wierstra, and T. Lillicrap · 2016
Closest in time.