Fetching the paper…
Reading the bibliography…
Model-based reinforcement learning approaches carry the promise of being data efficient.
Evolutionary principles in self-referential learning. on learning now to learn: The meta-meta-meta…-hook
J. Schmidhuber · 1987
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
R. S. Sutton and R. S · 1991
Earlier work this paper cites.
Robust and Optimal Control
K. Zhou, J. C. Doyle, and K. Glover · 1996
Earlier work this paper cites.
Fast learning for problem classes using a knowledge based network initialization
M. Hüsken and C. Goerick · 2000
Earlier work this paper cites.
Autonomous helicopter control using reinforcement learning policy search methods
J. A. Bagnell and J. G. Schneider · 2001
Earlier work this paper cites.
Using inaccurate models in reinforcement learning
P. Abbeel, M. Quigley, and A. Y. Ng · 2006
Earlier work this paper cites.
Policy gradient methods for robotics
J. Peters and S. Schaal · 2006
Earlier work this paper cites.
Local gaussian process regression for real time online model learning and control
D. Nguyen-Tuong, M. Seeger, and J. Peters · 2009
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
M. Deisenroth and C. E. Rasmussen · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
A survey on policy search for robotics
M. P. Deisenroth, G. Neumann, and J. Peters · 2013
Earlier work this paper cites.
Reinforcement learning in robust markov decision processes
S. H. Lim, H. Xu, and S. Mannor · 2013
Earlier work this paper cites.
Learning neural network policies with guided policy search under unknown dynamics
S. Levine and P. Abbeel · 2014
Earlier work this paper cites.
Trust Region Policy Optimization
J. Schulman, S. Levine, P. Moritz, M. I. Jordan, and P. Abbeel · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap et al · 2015
Earlier work this paper cites.
Deep learning helicopter dynamics models
A. Punjani and P. Abbeel · 2015
Cited alongside, same era.
From pixels to torques: Policy learning with deep dynamical models
N. Wahlström, T. B. Schön, and M. P. Deisenroth · 2015
Cited alongside, same era.
One-shot learning of manipulation skills with online dynamics adaptation and neural network priors
J. Fu, S. Levine, and P. Abbeel · 2015
Cited alongside, same era.
Deepmpc: Learning deep latent features for model predictive control
I. Lenz, R. A. Knepper, and A. Saxena · 2015
Cited alongside, same era.
Learning Continuous Control Policies by Stochastic Value Gradients
N. Heess et al · 2015
Cited alongside, same era.
Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
C. Finn, P. Abbeel, and S. Levine · 2017
Later among the works it cites.
Learning to reinforcement learn
J. X. Wang et al · 2017
Later among the works it cites.
Learning to learn: Meta-critic networks for sample efficient learning
F. Sung, L. Zhang, T. Xiang, T. M. Hospedales, and Y. Yang · 2017
Later among the works it cites.
Data-efficient reinforcement learning with probabilistic model predictive control
S. Kamthe and M. P. Deisenroth · 2017
Later among the works it cites.
Neural Network Dynamics for Model-Based Deep Reinforcement Learning with Model-Free Fine-Tuning
A. Nagabandi, G. Kahn, R. S. Fearing, and S. Levine · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Silver et al · 2016
Cited alongside, same era.
EPOpt: Learning Robust Neural Network Policies Using Model Ensembles
A. Rajeswaran, S. Ghotra, B. Ravindran, and S. Levine · 2016
Cited alongside, same era.
RL$ˆ2$: Fast Reinforcement Learning via Slow Reinforcement Learning
Y. Duan, J. Schulman, X. Chen, P. L. Bartlett, I. Sutskever, and P. Abbeel · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Cited alongside, same era.
Continuous deep Q-learning with model-based acceleration
S. Gu, T. Lillicrap, I. Sutskever, and S. Levine · 2016
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
M. Andrychowicz, M. Denil, S. G. Colmenarejo, M. W. Hoffman, D. Pfau, T. Schaul, and N. de Freitas · 2016
Cited alongside, same era.
One-shot learning with memory-augmented neural networks
A. Santoro, S. Bartunov, M. Botvinick, D. Wierstra, and T. Lillicrap · 2016
Cited alongside, same era.
Learning and Policy Search in Stochastic Dynamical Systems with Bayesian Neural Networks
S. Depeweg, F. Doshi-velez, and S. Udluft · 2017
Later among the works it cites.
Prediction and Control with Temporal Segment Models
N. Mishra, P. Abbeel, and I. Mordatch · 2017
Later among the works it cites.
Proximal Policy Optimization Algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Later among the works it cites.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Y. Wu, E. Mansimov, S. Liao, R. B. Grosse, and J. Ba · 2017
Later among the works it cites.
Temporal Difference Models: Model-Free Deep RL for Model-Based Control
V. Pong, S. Gu, M. Dalal, and S. Levine · 2018
Closest in time.
Model-Ensemble Trust-Region Policy Optimization
T. Kurutach, I. Clavera, Y. Duan, A. Tamar, and P. Abbeel · 2018
Closest in time.
A Simple Neural Attentive Meta-Learner
N. Mishra, M. Rohaninejad, X. Chen, and P. Abbeel · 2018
Closest in time.
Learning to adapt: Meta-learning for model-based control
I. Clavera, A. Nagabandi, R. S. Fearing, P. Abbeel, S. Levine, and C. Finn · 2018
Closest in time.
Model-Based Value Expansion for Efficient Model-Free Reinforcement Learning
V. Feinberg, A. Wan, I. Stoica, M. I. Jordan, J. E. Gonzalez, and S. Levine · 2018
Closest in time.
Optimization as a model for few-shot learning
S. Ravi and H. Larochelle · 2018
Closest in time.
Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models
K. Chua, R. Calandra, R. Mcallister, and S. Levine · 2019
Closest in time.