Fetching the paper…
Reading the bibliography…
Model-based meta-reinforcement learning (RL) methods have recently been shown to be a promising approach to improving the sample efficiency of RL in multi-task settings.
Dyna, an integrated architecture for learning, planning, and reacting
Richard S Sutton · 1991
Earlier work this paper cites.
Monte-Carlo planning in large POMDPs
David Silver and Joel Veness · 2010
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Model predictive path integral control using covariance variable importance sampling
Grady Williams, Andrew Aldrich, and Evangelos Theodorou · 2015
Earlier work this paper cites.
Learning to reinforcement learn
Jane X Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z Leibo, Remi Munos, Charles Blundell, Dharshan Kumaran, and Matt Botvinick · 2016
Earlier work this paper cites.
RL 2 : Fast reinforcement learning via slow reinforcement learning
Yan Duan, John Schulman, Xi Chen, Peter L Bartlett, Ilya Sutskever, and Pieter Abbeel · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Searching for activation functions
Prajit Ramachandran, Barret Zoph, and Quoc V Le · 2017
Earlier work this paper cites.
Continuous adaptation via meta-learning in nonstationary and competitive environments
Maruan Al-Shedivat, Trapit Bansal, Yuri Burda, Ilya Sutskever, Igor Mordatch, and Pieter Abbeel · 2018
Earlier work this paper cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine · 2018
Earlier work this paper cites.
Model-based reinforcement learning via meta-policy optimization
Ignasi Clavera, Jonas Rothfuss, John Schulman, Yasuhiro Fujita, Tamim Asfour, and Pieter Abbeel · 2018
Cited alongside, same era.
Model-based value expansion for efficient model-free reinforcement learning
Vladimir Feinberg, Alvin Wan, Ion Stoica, Michael I. Jordan, Joseph E. Gonzalez, and Sergey Levine · 2018
Cited alongside, same era.
Meta-learning and universality: Deep representations and gradient descent can approximate any learning algorithm
Chelsea Finn and Sergey Levine · 2018
Cited alongside, same era.
Meta-reinforcement learning of structured exploration strategies
Abhishek Gupta, Russell Mendonca, YuXuan Liu, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Later among the works it cites.
Guided meta-policy search
Russell Mendonca, Abhishek Gupta, Rosen Kralev, Pieter Abbeel, Sergey Levine, and Chelsea Finn · 2019
Later among the works it cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Kate Rakelly, Aurick Zhou, Chelsea Finn, Sergey Levine, and Deirdre Quillen · 2019
Later among the works it cites.
ProMP: Proximal meta-policy search
Jonas Rothfuss, Dennis Lee, Ignasi Clavera, Tamim Asfour, and Pieter Abbeel · 2019
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Chelsea Finn, and Sergey Levine · 2019
Later among the works it cites.
Learning dexterous in-hand manipulation
Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafal Józefowicz, Bob McGrew, Jakub Pachocki, Arthur Petron, Matthias Plappert, Glenn Powell, Alex Ray, Jonas Schneider, Szymon Sidor, Josh Tobin, Peter Welinder, Lilian Weng, and Wojciech Zaremba · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Algorithmic framework for model-based deep reinforcement learning with theoretical guarantees
Yuping Luo, Huazhe Xu, Yuanzhi Li, Yuandong Tian, Trevor Darrell, and Tengyu Ma · 2018
Cited alongside, same era.
A simple neural attentive meta-learner
Nikhil Mishra, Mostafa Rohaninejad, Xi Chen, and Pieter Abbeel · 2018
Cited alongside, same era.
Meta reinforcement learning with latent variable Gaussian processes
Steindór Sæmundsson, Katja Hofmann, and Marc Peter Deisenroth · 2018
Cited alongside, same era.
Some considerations on learning to explore via meta-reinforcement learning
Bradly C Stadie, Ge Yang, Rein Houthooft, Xi Chen, Yan Duan, Yuhuai Wu, Pieter Abbeel, and Ilya Sutskever · 2018
Cited alongside, same era.
Explicit explore-exploit algorithms in continuous state spaces
Mikael Henaff · 2019
Cited alongside, same era.
Meta reinforcement learning as task inference
Jan Humplik, Alexandre Galashov, Leonard Hasenclever, Pedro A Ortega, Yee Whye Teh, and Nicolas Heess · 2019
Cited alongside, same era.
Learning to adapt in dynamic, real-world environments via meta-reinforcement learning
Anusha Nagabandi, Ignasi Clavera, Simin Liu, Ronald S Fearing, Pieter Abbeel, Sergey Levine, and Chelsea Finn
Cited in the paper.
Closest in time.
Generalized hidden parameter MDPs transferable model-based RL in a handful of trials
Christian F Perez, Felipe Petroski Such, and Theofanis Karaletsos · 2020
Closest in time.
A game theoretic framework for model based reinforcement learning, 2020
Aravind Rajeswaran, Igor Mordatch, and Vikash Kumar · 2020
Closest in time.
Model-based policy optimization with unsupervised model adaptation
Jian Shen, Han Zhao, Weinan Zhang, and Yong Yu · 2020
Closest in time.
MOPO: model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Y Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Closest in time.
VariBAD: A very good method for Bayes-adaptive deep RL via meta-learning
Luisa Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze, Yarin Gal, Katja Hofmann, and Shimon Whiteson · 2020
Closest in time.