Fetching the paper…
Reading the bibliography…
Model-based reinforcement learning (MBRL) aims to learn model(s) of the environment dynamics that can predict the outcome of its actions.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S. Sutton · 1990
Earlier work this paper cites.
The three sigma rule
Friedrich Pukelsheim · 1994
Earlier work this paper cites.
Introduction to Reinforcement Learning
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Model Predictive Control
E.F. Camacho, C. Bordons, and C.B. Alba · 2004
Earlier work this paper cites.
Muhammad Burhan Hafez, Cornelius Weber, Matthias Kerzel, and Stefan Wermter · 2004
Earlier work this paper cites.
A survey of numerical methods for optimal control
Anvil V. Rao · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Chapter 3 - the cross-entropy method for optimization
Zdravko I. Botev, Dirk P. Kroese, Reuven Y. Rubinstein, and Pierre L’Ecuyer · 2013
Earlier work this paper cites.
Marc Peter Deisenroth, Gerhard Neumann, and Jan Peters, 2013
2013
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients, 2015
Nicolas Heess, Greg Wayne, David Silver, Timothy Lillicrap, Yuval Tassa, and Tom Erez · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles, 2016
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2016
Cited alongside, same era.
Aggressive deep driving: Model predictive control with a cnn cost model
Paul Drews, Brian Goldfain, Grady Williams, and Evangelos A. Theodorou · 2017
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Jacob Buckman, Danijar Hafner, George Tucker, Eugene Brevdo, and Honglak Lee · 2018
Later among the works it cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models, 2018
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, 2018
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Anusha Nagabandi, G. Kahn, Ronald S. Fearing, and S. Levine · 2018
Later among the works it cites.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and S. Levine · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Uncertainty-driven imagination for continuous deep reinforcement learning
Gabriel Kalweit and Joschka Boedecker · 2017
Cited alongside, same era.
Value prediction network
Junhyuk Oh, Satinder Singh, and Honglak Lee · 2017
Cited alongside, same era.
Proximal policy optimization algorithms, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Information theoretic model predictive control: Theory and applications to autonomous driving, 2017
Grady Williams, Paul Drews, Brian Goldfain, James M. Rehg, and Evangelos A. Theodorou · 2017
Cited alongside, same era.
Dream to control: Learning behaviors by latent imagination, 2020
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization, 2020
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Later among the works it cites.
Bridging imagination and reality for model-based deep reinforcement learning
Guangxiang Zhu, Minghao Zhang, Honglak Lee, and Chongjie Zhang · 2020
Later among the works it cites.
Temporal difference learning for model predictive control, 2022
Nicklas Hansen, Xiaolong Wang, and Hao Su · 2022
Closest in time.