Stabilizing off-policy q-learning via bootstrapping error reduction
Original
Aviral Kumar, Justin Fu, George Tucker, and Sergey Levine · 1906
Earlier work this paper cites.
Torchmeta: A Meta-Learning library for PyTorch, 2019
Original
Tristan Deleu, Tobias Würfl, Mandana Samiei, Joseph Paul Cohen, and Yoshua Bengio · 1909
Earlier work this paper cites.
Evolutionary principles in self-referential learning
Jurgen Schmidhuber · 1987
Earlier work this paper cites.
On the optimization of a synaptic learning rule
Samy Bengio, Yoshua Bengio, Jocelyn Cloutier, and Jan Gecsei · 1992
Earlier work this paper cites.
Learning to learn: Introduction and overview
Sebastian Thrun and Lorien Pratt · 1998
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
Jan Peters and Stefan Schaal · 2007
Earlier work this paper cites.
Adam: A method for stochastic optimization
Original
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Human-level concept learning through probabilistic program induction
B. M. Lake, R. Salakhutdinov, and J. B. Tenenbaum · 2015
Earlier work this paper cites.
Rl2: Fast reinforcement learning via slow reinforcement learning
Original
Yan Duan, John Schulman, Xi Chen, Peter L Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
Hypernetworks, 2016
David Ha, Andrew Dai, and Quoc V. Le · 2016
Earlier work this paper cites.
Matching networks for one shot learning
Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Koray Kavukcuoglu, and Daan Wierstra · 2016
Earlier work this paper cites.
Learning to reinforcement learn
Original
Jane X Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z Leibo, Remi Munos, Charles Blundell, Dharshan Kumaran, and Matt Botvinick · 2016
Earlier work this paper cites.
Meta-learning with temporal convolutions
Original
Nikhil Mishra, Mostafa Rohaninejad, Xi Chen, and Pieter Abbeel · 2017
Earlier work this paper cites.
Meta-learning and universality: Deep representations and gradient descent can approximate any learning algorithm
Chelsea Finn and Sergey Levine · 2018
Earlier work this paper cites.
Learning to Learn with Gradients
Chelsea B Finn · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Earlier work this paper cites.