Fetching the paper…
Reading the bibliography…
Model-based reinforcement learning (RL) is considered to be a promising approach to reduce the sample complexity that hinders model-free RL.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton · 1990
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Richard S Sutton · 1991
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Ronald J Williams and Jing Peng · 1991
Earlier work this paper cites.
Neural networks for control systems—a survey
K Jetal Hunt, D Sbarbaro, R Żbikowski, and Peter J Gawthrop · 1992
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Minimax differential dynamic programming: An application to robust biped walking
Jun Morimoto and Christopher G Atkeson · 2003
Earlier work this paper cites.
Lipschitz continuity of value functions in markovian decision processes
Karl Hinderer · 2005
Earlier work this paper cites.
Regal: A regularization based algorithm for reinforcement learning in weakly communicating mdps
Peter L Bartlett and Ambuj Tewari · 2009
Earlier work this paper cites.
Gp-bayesfilters: Bayesian filtering using gaussian process prediction and observation models
Jonathan Ko and Dieter Fox · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Model-based reinforcement learning with nearly tight exploration complexity bounds
István Szita and Csaba Szepesvári · 2010
Earlier work this paper cites.
Regret bounds for the adaptive control of linear quadratic systems
Yasin Abbasi-Yadkori and Csaba Szepesvári · 2011
Earlier work this paper cites.
Learning to control a low-cost manipulator using data-efficient reinforcement learning
Marc Peter Deisenroth, Carl Edward Rasmussen, and Dieter Fox · 2011
Earlier work this paper cites.
Learning stable nonlinear dynamical systems with gaussian mixture models
S Mohammad Khansari-Zadeh and Aude Billard · 2011
Earlier work this paper cites.
Elements of information theory
Thomas M Cover and Joy A Thomas · 2012
Earlier work this paper cites.
Dyna-style planning with linear function approximation and prioritized sweeping
Richard S Sutton, Csaba Szepesvári, Alborz Geramifard, and Michael P Bowling · 2012
Earlier work this paper cites.
Integrating a partial model into model free reinforcement learning
Aviv Tamar, Dotan Di Castro, and Ron Meir · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
A survey on policy search for robotics
Marc Peter Deisenroth, Gerhard Neumann, Jan Peters, et al · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Guided policy search
Sergey Levine and Vladlen Koltun · 2013
Earlier work this paper cites.
Adaptive step-size for policy gradient methods
Matteo Pirotta, Marcello Restelli, and Luca Bascetta · 2013
Earlier work this paper cites.
Learning neural network policies with guided policy search under unknown dynamics
Sergey Levine and Pieter Abbeel · 2014
Earlier work this paper cites.
Sample-based informationl-theoretic stochastic optimal control
Rudolf Lioutikov, Alexandros Paraschos, Jan Peters, and Gerhard Neumann · 2014
Cited alongside, same era.
On the chi square and higher-order chi distances for approximating f-divergences
Frank Nielsen and Richard Nock · 2014
Cited alongside, same era.
Model-based policy gradients with parameter-based exploration by least-squares conditional density estimation
Voot Tangkaratt, Syogo Mori, Tingting Zhao, Jun Morimoto, and Masashi Sugiyama · 2014
Cited alongside, same era.
Model-less feedback control of continuum manipulators in constrained environments
Michael C Yip and David B Camarillo · 2014
Cited alongside, same era.
Improved regret bounds for undiscounted continuous reinforcement learning
Kailasam Lakshmanan, Ronald Ortner, and Daniil Ryabko · 2015
Cited alongside, same era.
Learning model-based planning from scratch
Razvan Pascanu, Yujia Li, Oriol Vinyals, Nicolas Heess, Lars Buesing, Sebastien Racanière, David Reichert, Théophane Weber, Daan Wierstra, and Peter Battaglia · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Sébastien Racanière, Théophane Weber, David Reichert, Lars Buesing, Arthur Guez, Danilo Jimenez Rezende, Adrià Puigdomènech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li, et al · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Optimism-driven exploration for nonlinear systems
Teodor Mihai Moldovan, Sergey Levine, Michael I Jordan, and Pieter Abbeel · 2015
Cited alongside, same era.
Policy gradient in lipschitz markov decision processes
Matteo Pirotta, Marcello Restelli, and Luca Bascetta · 2015
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Cited alongside, same era.
Continuous deep q-learning with model-based acceleration
Shixiang Gu, Timothy Lillicrap, Ilya Sutskever, and Sergey Levine · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Cited alongside, same era.
Kavosh Asadi, Dipendra Misra, and Michael L Littman · 2018
Closest in time.
Finite-data performance guarantees for the output-feedback control of an unknown system
Ross Boczar, Nikolai Matni, and Benjamin Recht · 2018
Closest in time.
Sample-Efficient Reinforcement Learning with Stochastic Ensemble Value Expansion
J. Buckman, D. Hafner, G. Tucker, E. Brevdo, and H. Lee · 2018
Closest in time.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine · 2018
Closest in time.
Model-based reinforcement learning via meta-policy optimization
Ignasi Clavera, Jonas Rothfuss, John Schulman, Yasuhiro Fujita, Tamim Asfour, and Pieter Abbeel · 2018
Closest in time.
Regret bounds for robust adaptive control of the linear quadratic regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2018
Closest in time.
Model-based value estimation for efficient model-free reinforcement learning
Vladimir Feinberg, Alvin Wan, Ion Stoica, Michael I Jordan, Joseph E Gonzalez, and Sergey Levine · 2018
Closest in time.
Efficient bias-span-constrained exploration-exploitation in reinforcement learning
Ronan Fruit, Matteo Pirotta, Alessandro Lazaric, and Ronald Ortner · 2018
Closest in time.
David Ha and Jürgen Schmidhuber · 2018
Closest in time.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Closest in time.
Variance Reduction Methods for Sublinear Reinforcement Learning
S. Kakade, M. Wang, and L. F. Yang · 2018
Closest in time.
Model-ensemble trust-region policy optimization
Thanard Kurutach, Ignasi Clavera, Yan Duan, Aviv Tamar, and Pieter Abbeel · 2018
Closest in time.
Simple random search provides a competitive approach to reinforcement learning
Horia Mania, Aurelia Guy, and Benjamin Recht · 2018
Closest in time.
Zero-shot visual imitation
Deepak Pathak, Parsa Mahmoudieh, Guanghao Luo, Pulkit Agrawal, Dian Chen, Yide Shentu, Evan Shelhamer, Jitendra Malik, Alexei A Efros, and Trevor Darrell · 2018
Closest in time.
Graph networks as learnable physics engines for inference and control
Alvaro Sanchez-Gonzalez, Nicolas Heess, Jost Tobias Springenberg, Josh Merel, Martin Riedmiller, Raia Hadsell, and Peter Battaglia · 2018
Closest in time.
The bottleneck simulator: A model-based deep reinforcement learning approach
Iulian Vlad Serban, Chinnadhurai Sankar, Michael Pieper, Joelle Pineau, and Yoshua Bengio · 2018
Closest in time.
Learning without mixing: Towards a sharp analysis of linear system identification
Max Simchowitz, Horia Mania, Stephen Tu, Michael I Jordan, and Benjamin Recht · 2018
Closest in time.
Aravind Srinivas, Allan Jabri, Pieter Abbeel, Sergey Levine, and Chelsea Finn · 2018
Closest in time.
Wen Sun, Geoffrey J Gordon, Byron Boots, and J Andrew Bagnell · 2018
Closest in time.
Provable model-based nonlinear bandit and reinforcement learning: Shelve optimism, embrace virtual curvature, 2021
Kefan Dong, Jiaqi Yang, and Tengyu Ma · 2021
Closest in time.