Fetching the paper…
Reading the bibliography…
Learning to learn has emerged as an important direction for achieving artificial intelligence.
Reminiscence and rote learning
Ward, Lewis B · 1937
Earlier work this paper cites.
The formation of learning sets
Harlow, Harry F · 1949
Earlier work this paper cites.
Evolutionary Principles in Self-Referential Learning. On Learning how to Learn: The Meta-Meta-Meta…-Hook
Schmidhuber, Jurgen · 1987
Earlier work this paper cites.
A layered network model of associative learning: learning to learn and configuration
Kehoe, E James · 1988
Earlier work this paper cites.
Learning a synaptic learning rule
Bengio, Yoshua, Bengio, Samy, and Cloutier, Jocelyn · 1990
Earlier work this paper cites.
On the optimization of a synaptic learning rule
Bengio, Yoshua, Bengio, Samy, Cloutier, Jocelyn, and Gecsei, Jan · 1992
Earlier work this paper cites.
Meta-neural networks that learn by learning
Naik, Devang K and Mammone, RJ · 1992
Earlier work this paper cites.
Adapting bias by gradient descent: An incremental version of delta-bar-delta
Sutton, Richard S · 1992
Earlier work this paper cites.
On the search for new learning rules for ANNs
Bengio, S., Bengio, Y., and Cloutier, J · 1995
Earlier work this paper cites.
Learning to learn
Thrun, Sebastian and Pratt, Lorien · 1998
Earlier work this paper cites.
An incremental gradient (-projection) method with momentum term and adaptive stepsize rule
Tseng, Paul · 1998
Earlier work this paper cites.
Evolution and design of distributed learning rules
Runarsson, Thomas Philip and Jonsson, Magnus Thor · 2000
Cited alongside, same era.
Learning to learn using gradient descent
Hochreiter, Sepp, Younger, A Steven, and Conwell, Peter R · 2001
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, Tijmen and Hinton, Geoffrey · 2012
Cited alongside, same era.
Optimization test functions and datasets, 2013
Surjanovic, Sonja and Bingham, Derek · 2013
Cited alongside, same era.
On the properties of neural machine translation: Encoder-decoder approaches
Cho, Kyunghyun, Van Merriënboer, Bart, Bahdanau, Dzmitry, and Bengio, Yoshua · 2014
Rl 2 : Fast reinforcement learning via slow reinforcement learning
Duan, Yan, Schulman, John, Chen, Xi, Bartlett, Peter, Sutskever, Ilya, and Abbeel, Pieter · 2016
Later among the works it cites.
Identity mappings in deep residual networks
He, Kaiming, Zhang, Xiangyu, Ren, Shaoqing, and Sun, Jian · 2016
Later among the works it cites.
Building machines that learn and think like people
Lake, Brenden M, Ullman, Tomer D, Tenenbaum, Joshua B, and Gershman, Samuel J · 2016
Later among the works it cites.
Meta-learning with memory-augmented neural networks
Santoro, ADAM, Bartunov, Sergey, Botvinick, Matthew, Wierstra, Daan, and Lillicrap, Timothy · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
Silver, David, Huang, Aja, Maddison, Chris J, Guez, Arthur, Sifre, Laurent, Van Den Driessche, George, Schrittwieser, Julian, Antonoglou, Ioannis, Panneershelvam, Veda, Lanctot, Marc, et al · 2016
Later among the works it cites.
Rethinking the inception architecture for computer vision
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2015
Cited alongside, same era.
RMSprop loses to SMORMS3 - beware the epsilon!, 2015
Funk, Simon · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2015
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
Andrychowicz, Marcin, Denil, Misha, Gomez, Sergio, Hoffman, Matthew W, Pfau, David, Schaul, Tom, Shillingford, Brendan, and de Freitas, Nando · 2016
Cited alongside, same era.
Learning to learn for global optimization of black box functions
Chen, Yutian, Hoffman, Matthew W., Colmenarejo, Sergio Gomez, Denil, Misha, Lillicrap, Timothy P., and de Freitas, Nando · 2016
Cited alongside, same era.
A method of solving a convex programming problem with convergence rate o (1/k2)
Nesterov, Yurii
Cited in the paper.
Szegedy, Christian, Vanhoucke, Vincent, Ioffe, Sergey, Shlens, Jon, and Wojna, Zbigniew · 2016
Later among the works it cites.
Learning to reinforcement learn
Wang, Jane X., Kurth-Nelson, Zeb, Tirumala, Dhruva, Soyer, Hubert, Leibo, Joel Z., Munos, Rémi, Blundell, Charles, Kumaran, Dharshan, and Botvinick, Matt · 2016
Later among the works it cites.
Learning to optimize
Li, SKe and Malik, Jitendra · 2017
Closest in time.
Optimization as a model for few-shot learning
Ravi, Sachin and Larochelle, Hugo · 2017
Closest in time.
Neural architecture search with reinforcement learning
Zoph, Barret and Le, Quoc V · 2017
Closest in time.