Fetching the paper…
Reading the bibliography…
Provably efficient Model-Based Reinforcement Learning (MBRL) based on optimism or posterior sampling (PSRL) is ensured to attain the global optimality asymptotically by introducing the complexity measure of the model.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
Dual effect, certainty equivalence, and separation in stochastic control
Yaakov Bar-Shalom and Edison Tse · 1974
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton · 1990
Earlier work this paper cites.
Bayesian q-learning
Richard Dearden, Nir Friedman, and Stuart Russell · 1998
Earlier work this paper cites.
Approximate planning in large pomdps via reusable trajectories
Michael Kearns, Yishay Mansour, and Andrew Ng · 1999
Earlier work this paper cites.
Model predictive control: past, present and future
Manfred Morari and Jay H Lee · 1999
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
A bayesian framework for reinforcement learning
Malcolm Strens · 2000
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
A theoretical analysis of model-based interval estimation
Alexander L Strehl and Michael L Littman · 2005
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Regret bounds for the adaptive control of linear quadratic systems
Yasin Abbasi-Yadkori and Csaba Szepesvári · 2011
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Marc Deisenroth and Carl E Rasmussen · 2011
Earlier work this paper cites.
Agnostic system identification for model-based reinforcement learning
Stephane Ross and J Andrew Bagnell · 2012
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Daniel Russo and Benjamin Van Roy · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Learning neural network policies with guided policy search under unknown dynamics
Sergey Levine and Pieter Abbeel · 2014
Earlier work this paper cites.
Model-based reinforcement learning and the eluder dimension
Ian Osband and Benjamin Van Roy · 2014
Earlier work this paper cites.
Learning to optimize via posterior sampling
Daniel Russo and Benjamin Van Roy · 2014
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Timothy Lillicrap, Tom Erez, and Yuval Tassa · 2015
Earlier work this paper cites.
Policy gradient in lipschitz markov decision processes
Matteo Pirotta, Marcello Restelli, and Luca Bascetta · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Dual control for approximate bayesian reinforcement learning
Edgar D Klenske and Philipp Hennig · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Cited alongside, same era.
Guided policy search as approximate mirror descent
William Montgomery and Sergey Levine · 2016
Cited alongside, same era.
Contextual decision processes with low bellman rank are pac-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Cited alongside, same era.
Ensemble sampling
Xiuyuan Lu and Benjamin Van Roy · 2017
Reinforcement learning: Theory and algorithms
Alekh Agarwal, Nan Jiang, Sham M Kakade, and Wen Sun · 2019
Later among the works it cites.
Online learning in kernelized markov decision processes
Sayak Ray Chowdhury and Aditya Gopalan · 2019
Later among the works it cites.
Dream to control: Learning behaviors by latent imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi · 2019
Later among the works it cites.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Later among the works it cites.
Certainty equivalence is efficient for linear quadratic control
Horia Mania, Stephen Tu, and Benjamin Recht · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Why is posterior sampling better than optimism for reinforcement learning?
Ian Osband and Benjamin Van Roy · 2017
Cited alongside, same era.
A tutorial on thompson sampling
Daniel Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, and Zheng Wen · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Maximum a posteriori policy optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, and Martin Riedmiller · 2018
Cited alongside, same era.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Jacob Buckman, Danijar Hafner, George Tucker, Eugene Brevdo, and Honglak Lee · 2018
Cited alongside, same era.
Bayesian methods for efficient reinforcement learning in tabular problems, 2019
Efstratios Markou and Carl E. Rasmussen · 2019
Later among the works it cites.
Deep exploration via randomized value functions
Ian Osband, Benjamin Van Roy, Daniel J Russo, Zheng Wen, et al · 2019
Later among the works it cites.
Model-based rl in contextual decision processes: Pac bounds and exponential improvements over model-free approaches
Wen Sun, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2019
Later among the works it cites.
Exploring model-based planning with policy networks
Tingwu Wang and Jimmy Ba · 2019
Later among the works it cites.
Sample-optimal parametric q-learning using linearly additive features
Lin Yang and Mengdi Wang · 2019
Later among the works it cites.
Sample complexity of estimating the policy gradient for nearly deterministic dynamical systems
Osbert Bastani · 2020
Later among the works it cites.
Model-augmented actor-critic: Backpropagating through paths
Ignasi Clavera, Violet Fu, and Pieter Abbeel · 2020
Later among the works it cites.
Efficient model-based reinforcement learning through optimistic policy search and planning
Sebastian Curi, Felix Berkenkamp, and Andreas Krause · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Later among the works it cites.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
Lin Yang and Mengdi Wang · 2020
Later among the works it cites.
Kefan Dong, Jiaqi Yang, and Tengyu Ma · 2021
Later among the works it cites.
Bilinear classes: A structural framework for provable generalization in rl
Simon Du, Sham Kakade, Jason Lee, Shachar Lovett, Gaurav Mahajan, Wen Sun, and Ruosong Wang · 2021
Later among the works it cites.
Optimism is all you need: Model-based imitation learning from observation alone
Rahul Kidambi, Jonathan Chang, and Wen Sun · 2021
Later among the works it cites.
Eluder dimension and generalized rank
Gene Li, Pritish Kamath, Dylan J Foster, and Nathan Srebro · 2021
Later among the works it cites.
Model-free representation learning and exploration in low-rank mdps
Aditya Modi, Jinglin Chen, Akshay Krishnamurthy, Nan Jiang, and Alekh Agarwal · 2021
Later among the works it cites.
Do differentiable simulators give better policy gradients?
HJ Suh, Max Simchowitz, Kaiqing Zhang, and Russ Tedrake · 2022
Closest in time.
Accelerated policy learning with parallel differentiable simulation
Jie Xu, Viktor Makoviychuk, Yashraj Narang, Fabio Ramos, Wojciech Matusik, Animesh Garg, and Miles Macklin · 2022
Closest in time.