Fetching the paper…
Reading the bibliography…
A promising way to improve the sample efficiency of reinforcement learning is model-based methods, in which many explorations and evaluations can happen in the learned models to save real-world samples.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
R. S. Sutton · 1990
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Chapter 3 - the cross-entropy method for optimization
Z. I. Botev, D. P. Kroese, R. Y. Rubinstein, and P. L’Ecuyer · 2013
Earlier work this paper cites.
Model Predictive Control
E. Camacho and C. Alba · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. A. Riedmiller, A. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. I. Jordan, and P. Moritz · 2015
Earlier work this paper cites.
Path integral networks: End-to-end differentiable optimal control
M. Okada, L. Rigazio, and T. Aoshima · 2017
Earlier work this paper cites.
Towards a simple approach to multi-step model-based reinforcement learning
K. Asadi, E. Cater, D. Misra, and M. L. Littman · 2018
Earlier work this paper cites.
Multi-step reinforcement learning: A unifying algorithm
K. D. Asis, J. F. Hernandez-Garcia, G. Z. Holland, and R. S. Sutton · 2018
Earlier work this paper cites.
Combining model-based and model-free RL via multi-step control variates
T. Che, Y. Lu, G. Tucker, S. Bhupatiraju, S. Gu, S. Levine, and Y. Bengio · 2018
Earlier work this paper cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
K. Chua, R. Calandra, R. McAllister, and S. Levine · 2018
Cited alongside, same era.
Model-based value estimation for efficient model-free reinforcement learning
V. Feinberg, A. Wan, I. Stoica, M. I. Jordan, J. E. Gonzalez, and S. Levine · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
A. Nagabandi, G. Kahn, R. S. Fearing, and S. Levine · 2018
Cited alongside, same era.
Universal planning networks: Learning generalizable representations for visuomotor control
A. Srinivas, A. Jabri, P. Abbeel, S. Levine, and C. Finn · 2018
Cited alongside, same era.
Algorithmic framework for model-based deep reinforcement learning with theoretical guarantees
Y. Luo, H. Xu, Y. Li, Y. Tian, T. Darrell, and T. Ma · 2019
Later among the works it cites.
Virtual-taobao: Virtualizing real-world online retail environment for reinforcement learning
J. Shi, Y. Yu, Q. Da, S. Chen, and A. Zeng · 2019
Later among the works it cites.
Model-augmented actor-critic: Backpropagating through paths
I. Clavera, Y. Fu, and P. Abbeel · 2020
Later among the works it cites.
Bidirectional model-based policy optimization
H. Lai, J. Shen, W. Zhang, and Y. Yu · 2020
Later among the works it cites.
Trust the model when it is confident: Masked model-based actor-critic
F. Pan, J. He, D. Tu, and Q. He · 2020
Later among the works it cites.
Partially observable environment estimation with uplift inference for reinforcement learning based recommendation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards sample efficient reinforcement learning
Y. Yu · 2018
Cited alongside, same era.
Combating the compounding-error problem with a multi-step model
K. Asadi, D. Misra, S. Kim, and M. L. Littman · 2019
Cited alongside, same era.
When to trust your model: Model-based policy optimization
M. Janner, J. Fu, M. Zhang, and S. Levine · 2019
Cited alongside, same era.
Differentiable algorithm networks for composable robot learning
P. Karkus, X. Ma, D. Hsu, L. P. Kaelbling, W. S. Lee, and T. Lozano-Pérez · 2019
Cited alongside, same era.
Modeling the long term future in model-based reinforcement learning
N. R. Ke, A. Singh, A. Touati, A. Goyal, Y. Bengio, D. Parikh, and D. Batra · 2019
Cited alongside, same era.
W. Shang, Q. Li, Z. Qin, Y. Yu, Y. Meng, and J. Ye · 2021
Later among the works it cites.
Error bounds of imitating policies and environments for reinforcement learning
T. Xu, Z. Li, and Y. Yu · 2021
Later among the works it cites.
Adversarial counterfactual environment model learning
X.-H. Chen, Y. Yu, Z.-M. Zhu, Z. Yu, Z. Chen, C. Wang, Y. Wu, H. Wu, R.-J. Qin, R. Ding, and F. Huang · 2022
Closest in time.
Generative planning for temporally coordinated exploration in reinforcement learning
H. Zhang, W. Xu, and H. Yu · 2022
Closest in time.
Offline reinforcement learning with causal structured world models
Z.-M. Zhu, X.-H. Chen, H.-L. Tian, K. Zhang, and Y. Yu · 2022
Closest in time.