Fetching the paper…
Reading the bibliography…
Significant progress has been made in the area of model-based reinforcement learning.
Planning by incremental dynamic programming
R. S. Sutton · 1991
Earlier work this paper cites.
Learning reactive admittance control
V. Gullapalli, R. A. Grupen, and A. G. Barto · 1992
Earlier work this paper cites.
Reinforcement learning: A survey
L. P. Kaelbling, M. L. Littman, and A. W. Moore · 1996
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
D. Precup, R. S. Sutton, and S. P. Singh · 2000
Earlier work this paper cites.
Off-policy temporal difference learning with function approximation
D. Precup, R. S. Sutton, and S. Dasgupta · 2001
Earlier work this paper cites.
Neural reinforcement learning controllers for a real robot application
R. Hafner and M. Riedmiller · 2007
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
M. Deisenroth and C. E. Rasmussen · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
B. Recht, C. Re, S. Wright, and F. Niu · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Synthesis and stabilization of complex behaviors through online trajectory optimization
Y. Tassa, T. Erez, and E. Todorov · 2012
Earlier work this paper cites.
Large scale distributed deep networks
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, M. aurelio Ranzato, A. Senior, P. Tucker, K. Yang, Q. V. Le, and A. Y. Ng · 2012
Earlier work this paper cites.
Guided policy search
S. Levine and V. Koltun · 2013
Earlier work this paper cites.
Sample-based informationl-theoretic stochastic optimal control
R. Lioutikov, A. Paraschos, J. Peters, and G. Neumann · 2014
Cited alongside, same era.
Learning contact-rich manipulation skills with guided policy search
S. Levine, N. Wagener, and P. Abbeel · 2015
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
N. Heess, G. Wayne, D. Silver, T. Lillicrap, T. Erez, and Y. Tassa · 2015
Cited alongside, same era.
Massively parallel methods for deep reinforcement learning
A. Nair, P. Srinivasan, S. Blackwell, C. Alcicek, R. Fearon, A. D. Maria, V. Panneershelvam, M. Suleyman, C. Beattie, S. Petersen, S. Legg, V. Mnih, K. Kavukcuoglu, and D. Silver · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Emergence of locomotion behaviours in rich environments
N. Heess, D. TB, S. Sriram, J. Lemmon, J. Merel, G. Wayne, Y. Tassa, T. Erez, Z. Wang, S. M. A. Eslami, M. A. Riedmiller, and D. Silver · 2017
Later among the works it cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Later among the works it cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
K. Chua, R. Calandra, R. McAllister, and S. Levine · 2018
Later among the works it cites.
Model-based reinforcement learning via meta-policy optimization
I. Clavera, J. Rothfuss, J. Schulman, Y. Fujita, T. Asfour, and P. Abbeel · 2018
Later among the works it cites.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimism-driven exploration for nonlinear systems
T. M. Moldovan, S. Levine, M. I. Jordan, and P. Abbeel · 2015
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Cited alongside, same era.
GA3C: gpu-based A3C for deep reinforcement learning
M. Babaeizadeh, I. Frosio, S. Tyree, J. Clemons, and J. Kautz · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Deep reinforcement learning for robotic manipulation
S. Gu, E. Holly, T. P. Lillicrap, and S. Levine · 2016
Cited alongside, same era.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
A. Nagabandi, G. Kahn, R. S. Fearing, and S. Levine · 2017
Cited alongside, same era.
J. Buckman, D. Hafner, G. Tucker, E. Brevdo, and H. Lee · 2018
Later among the works it cites.
Model-ensemble trust-region policy optimization
T. Kurutach, I. Clavera, Y. Duan, A. Tamar, and P. Abbeel · 2018
Later among the works it cites.
IMPALA: scalable distributed deep-rl with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, S. Legg, and K. Kavukcuoglu · 2018
Later among the works it cites.
Accelerated methods for deep reinforcement learning
A. Stooke and P. Abbeel · 2018
Later among the works it cites.
Learning to adapt: Meta-learning for model-based control
I. Clavera, A. Nagabandi, R. S. Fearing, P. Abbeel, S. Levine, and C. Finn · 2018
Later among the works it cites.
When to trust your model: Model-based policy optimization
M. Janner, J. Fu, M. Zhang, and S. Levine · 2019
Closest in time.
Benchmarking model-based reinforcement learning
T. Wang, X. Bao, I. Clavera, J. Hoang, Y. Wen, E. Langlois, S. Zhang, G. Zhang, P. Abbeel, and J. Ba · 2019
Closest in time.
Algorithmic framework for model-based deep reinforcement learning with theoretical guarantees
Y. Luo, H. Xu, Y. Li, Y. Tian, T. Darrell, and T. Ma · 2019
Closest in time.