Fetching the paper…
Reading the bibliography…
Credit assignment in Meta-reinforcement learning (Meta-RL) is still poorly understood.
Evolutionary principles in self-referential learning. On learning how to learn: The meta-meta-… hook
Juergen Schmidhuber · 1987
Earlier work this paper cites.
Shifting Inductive Bias with Success-Story Algorithm, Adaptive Levin Search, and Incremental Self-Improvement
Jürgen Schmidhuber, Jieyu Zhao, and Marco Wiering · 1997
Earlier work this paper cites.
Learning to learn
Sebastian Thrun and Lorien Pratt · 1998
Earlier work this paper cites.
Fast learning for problem classes using knowledge based network initialization
Michael Hüsken and Christian Goerick · 2000
Earlier work this paper cites.
Policy Gradient Methods for Reinforcement Learning with Function Approximation
Richard S. Sutton, David Mcallester, Satinder Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Infinite-Horizon Policy-Gradient Estimation
Jonathan Baxter and Peter L Bartlett · 2001
Earlier work this paper cites.
Learning To Learn Using Gradient Descent
Sepp Hochreiter, A. Steven Younger, and Peter R. Conwell · 2001
Earlier work this paper cites.
Approximately Optimal Approximate Reinforcement Learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Policy Gradient Methods for Robotics
Jan Peters and Stefan Schaal · 2006
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gómez Colmenarejo, Matthew W Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando De Freitas · 2016
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
RL$ˆ2$: Fast Reinforcement Learning via Slow Reinforcement Learning
Yan Duan, John Schulman, Xi Chen, Peter L. Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Cited alongside, same era.
Approximate Newton Methods for Policy Search in Markov Decision Processes
Thomas Furmston, Guy Lever, David Barber, and Joelle Pineau · 2016
Cited alongside, same era.
Meta-Learning with Memory-Augmented Neural Networks
Adam Santoro, Sergey Bartunov, Matthew Botvinick, Daan Wierstra, Timothy Lillicrap, and Google Deepmind · 2016
Cited alongside, same era.
Constrained Policy Optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Continuous Adaptation via Meta-Learning in Nonstationary and Competitive Environments
Maruan Al-Shedivat, Trapit Bansal, Umass Amherst, Yura Burda, Openai Ilya, Sutskever Openai, Igor Mordatch Openai, and Pieter Abbeel · 2018
Closest in time.
Ferran Alet, Tomás Lozano-Pérez, and Leslie P. Kaelbling · 2018
Closest in time.
Model-Based Reinforcement Learning via Meta-Policy Optimization
Ignasi Clavera, Jonas Rothfuss, John Schulman, Yasuhiro Fujita, Tamim Asfour, and Pieter Abbeel · 2018
Closest in time.
DiCE: The Infinitely Differentiable Monte Carlo Estimator
Jakob Foerster, Gregory Farquhar, Maruan Al-Shedivat, Tim Rocktäschel, Eric P Xing, and Shimon Whiteson · 2018
Closest in time.
Meta Learning Shared Hierarchies
Kevin Frans, Jonathan Ho, Xi Chen, Pieter Abbeel, and John Schulman · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning to Learn without Gradient Descent by Gradient Descent
Yutian Chen, Matthew W Hoffman, Sergio Gómez Colmenarejo, Misha Denil, Timothy P Lillicrap, Matt Botvinick, and Nando De Freitas · 2017
Cited alongside, same era.
Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Learning to Generalize: Meta-Learning for Domain Generalization
Da Li, Yongxin Yang, Yi-Zhe Song, and Timothy M Hospedales · 2017
Cited alongside, same era.
Optimization as a Model for Few-Shot Learning
Sachin Ravi and Hugo Larochelle · 2017
Cited alongside, same era.
Proximal Policy Optimization Algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov Openai · 2017
Cited alongside, same era.
Learning to Learn: Meta-Critic Networks for Sample Efficient Learning
Flood Sung, Li Zhang, Tao Xiang, Timothy Hospedales, and Yongxin Yang · 2017
Cited alongside, same era.
Unsupervised Meta-Learning for Reinforcement Learning
Abhishek Gupta, Benjamin Eysenbach, Chelsea Finn, and Sergey Levine
Cited in the paper.
Closest in time.
Differentiable plasticity: training plastic neural networks with backpropagation
Thomas Miconi, Jeff Clune, and Kenneth O. Stanley · 2018
Closest in time.
A Simple Neural Attentive Meta-Learner
Nikhil Mishra, Mostafa Rohaninejad, Xi Chen, and Pieter Abbeel · 2018
Closest in time.
On First-Order Meta-Learning Algorithms
Alex Nichol, Joshua Achiam, and John Schulman · 2018
Closest in time.
Meta Reinforcement Learning with Latent Variable Gaussian Processes
Steindór Saemundsson, Katja Hofmann, and Marc Peter Deisenroth · 2018
Closest in time.
Some Considerations on Learning to Explore via Meta-Reinforcement Learning
Bradly C Stadie, Ge Yang, Rein Houthooft, Xi Chen, Yan Duan, Yuhuai Wu, Pieter Abbeel, and Ilya Sutskever · 2018
Closest in time.
Meta-Gradient Reinforcement Learning
Zhongwen Xu, Hado van Hasselt, and David Silver · 2018
Closest in time.