Fetching the paper…
Reading the bibliography…
Model-Based Reinforcement Learning involves learning a \textit{dynamics model} from data, and then using this model to optimise behaviour, most often with an online \textit{planner}.
Estimation of inertial parameters of manipulator loads and links
C. G. Atkeson, C. H. An, and J. M. Hollerbach · 1986
Earlier work this paper cites.
Model predictive control: Theory and practice—a survey
C. E. Garcia, D. M. Prett, and M. Morari · 1989
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
R. S. Sutton · 1991
Earlier work this paper cites.
Neural networks for control
W. T. Miller, P. J. Werbos, and R. S. Sutton · 1995
Earlier work this paper cites.
Scalable techniques from nonparametric statistics for real time robot learning
S. Schaal, C. G. Atkeson, and S. Vijayakumar · 2002
Earlier work this paper cites.
Learning vehicular dynamics, with application to modeling helicopters
P. Abbeel, V. Ganapathi, and A. Ng · 2005
Earlier work this paper cites.
Local online support vector regression for learning control
Y. Choi, S.-Y. Cheong, and N. Schweighofer · 2007
Earlier work this paper cites.
Model learning with local gaussian process regression
D. Nguyen-Tuong, M. Seeger, and J. Peters · 2009
Earlier work this paper cites.
Learning and reproduction of gestures by imitation
S. Calinon, F. D’halluin, E. L. Sauser, D. G. Caldwell, and A. G. Billard · 2010
Earlier work this paper cites.
Using model knowledge for learning inverse dynamics
D. Nguyen-Tuong and J. Peters · 2010
Earlier work this paper cites.
Model learning for robot control: a survey
D. Nguyen-Tuong and J. Peters · 2011
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
M. Deisenroth and C. E. Rasmussen · 2011
Earlier work this paper cites.
Nonlinear model predictive control
F. Allgöwer and A. Zheng · 2012
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
Improving multi-step prediction of learned time series models
A. Venkatraman, M. Hebert, and J. A. Bagnell · 2015
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
N. Heess, G. Wayne, D. Silver, T. Lillicrap, Y. Tassa, and T. Erez · 2015
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, and W. Zaremba · 2017
Earlier work this paper cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
K. Chua, R. Calandra, R. McAllister, and S. Levine · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. van Hoof, and D. Meger · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Cited alongside, same era.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
A. Nagabandi, G. Kahn, R. S. Fearing, and S. Levine · 2018
Cited alongside, same era.
Mopo: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Zou, S. Levine, C. Finn, and T. Ma · 2020
Later among the works it cites.
A game theoretic framework for model based reinforcement learning
A. Rajeswaran, I. Mordatch, and V. Kumar · 2020
Later among the works it cites.
Morel: Model-based offline reinforcement learning
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Later among the works it cites.
Deep dynamics models for learning dexterous manipulation
A. Nagabandi, K. Konolige, S. Levine, and V. Kumar · 2020
Later among the works it cites.
Imagined value gradients: Model-based policy optimization with tranferable latent dynamics models
A. Byravan, J. T. Springenberg, A. Abdolmaleki, R. Hafner, M. Neunert, T. Lampe, N. Siegel, N. Heess, and M. Riedmiller · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Buckman, D. Hafner, G. Tucker, E. Brevdo, and H. Lee · 2018
Cited alongside, same era.
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. d. L. Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, et al · 2018
Cited alongside, same era.
Model-based value estimation for efficient model-free reinforcement learning
V. Feinberg, A. Wan, I. Stoica, M. I. Jordan, J. E. Gonzalez, and S. Levine · 2018
Cited alongside, same era.
Model-ensemble trust-region policy optimization
T. Kurutach, I. Clavera, Y. Duan, A. Tamar, and P. Abbeel · 2018
Cited alongside, same era.
D. Ha and J. Schmidhuber · 2018
Cited alongside, same era.
Graph networks as learnable physics engines for inference and control
A. Sanchez-Gonzalez, N. Heess, J. T. Springenberg, J. Merel, M. Riedmiller, R. Hadsell, and P. Battaglia · 2018
Cited alongside, same era.
Maximum a posteriori policy optimisation
A. Abdolmaleki, J. T. Springenberg, Y. Tassa, R. Munos, N. Heess, and M. Riedmiller · 2018
Cited alongside, same era.
Learning to simulate complex physics with graph networks
A. Sanchez-Gonzalez, J. Godwin, T. Pfaff, R. Ying, J. Leskovec, and P. Battaglia · 2020
Later among the works it cites.
Learning mesh-based simulation with graph networks
T. Pfaff, M. Fortunato, A. Sanchez-Gonzalez, and P. W. Battaglia · 2020
Later among the works it cites.
A differentiable newton euler algorithm for multi-body model learning
M. Lutter, J. Silberbauer, J. Watson, and J. Peters · 2020
Later among the works it cites.
Mastering atari with discrete world models
D. Hafner, T. Lillicrap, M. Norouzi, and J. Ba · 2020
Later among the works it cites.
Learning accurate long-term dynamics for model-based reinforcement learning
N. O. Lambert, A. Wilcox, H. Zhang, K. S. Pister, and R. Calandra · 2020
Later among the works it cites.
Sample-efficient cross-entropy method for real-time planning
C. Pinneri, S. Sawant, S. Blaes, J. Achterhold, J. Stueckler, M. Rolinek, and G. Martius · 2020
Later among the works it cites.
Objective mismatch in model-based reinforcement learning
N. Lambert, B. Amos, O. Yadan, and R. Calandra · 2020
Later among the works it cites.
Challenges of real-world reinforcement learning: definitions, benchmarks and analysis
G. Dulac-Arnold, N. Levine, D. J. Mankowitz, J. Li, C. Paduraru, S. Gowal, and T. Hester · 2021
Closest in time.
Model-based offline planning
A. Argenson and G. Dulac-Arnold · 2021
Closest in time.
Model-based meta-reinforcement learning for flight with suspended payloads
S. Belkhale, R. Li, G. Kahn, R. McAllister, R. Calandra, and S. Levine · 2021
Closest in time.
Differentiable physics models for real-world offline model-based reinforcement learning
M. Lutter, J. Silberbauer, J. Watson, and J. Peters · 2021
Closest in time.
Model predictive actor-critic: Accelerating robot skill acquisition with deep reinforcement learning
A. S. Morgan, D. Nandha, G. Chalvatzaki, C. D’Eramo, A. M. Dollar, and J. Peters · 2021
Closest in time.
On the model-based stochastic value gradient for continuous reinforcement learning
B. Amos, S. Stanton, D. Yarats, and A. G. Wilson · 2021
Closest in time.
Stochastic control through approximate bayesian input inference
J. Watson, H. Abdulsamad, R. Findeisen, and J. Peters · 2021
Closest in time.