Fetching the paper…
Reading the bibliography…
We propose a general framework to design posterior sampling methods for model-based RL.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
The equivalence of two extremum problems
Jack Kiefer and Jacob Wolfowitz · 1960
Earlier work this paper cites.
Banach-Mazur distances and finite-dimensional operator ideals , volume 38
Nicole Tomczak-Jaegermann · 1989
Earlier work this paper cites.
Exploiting structure in policy construction
Craig Boutilier, Richard Dearden, Moisés Goldszmidt, et al · 1995
Earlier work this paper cites.
Neuro-Dynamic Programming
Dimitri P. Bertsekas and John N. Tsitsiklis · 1996
Earlier work this paper cites.
Empirical Processes in M-estimation , volume 6
Sara van de Geer · 2000
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
On the generalization ability of on-line learning algorithms
Nicolo Cesa-Bianchi, Alex Conconi, and Claudio Gentile · 2004
Earlier work this paper cites.
From ϵ \epsilon -entropy to kl-entropy: Analysis of minimum information complexity density estimation
Tong Zhang · 2006
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Peter Auer, Thomas Jaksch, and Ronald Ortner · 2008
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
M.L. Puterman · 2014
Earlier work this paper cites.
Contextual markov decision processes
Assaf Hallak, Dotan Di Castro, and Shie Mannor · 2015
Earlier work this paper cites.
f f -divergence inequalities
Igal Sason and Sergio Verdú · 2016
Earlier work this paper cites.
Posterior sampling for reinforcement learning: worst-case regret bounds
Shipra Agrawal and Randy Jia · 2017
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Earlier work this paper cites.
Contextual decision processes with low bellman rank are pac-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Cited alongside, same era.
Ensemble sampling
Xiuyuan Lu and Benjamin Van Roy · 2017
Cited alongside, same era.
A tutorial on thompson sampling
Daniel Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, and Zheng Wen · 2017
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine · 2018
Cited alongside, same era.
David Ha and Jürgen Schmidhuber · 2018
Cited alongside, same era.
Information theoretic regret bounds for online nonlinear control
Sham M. Kakade, Akshay Krishnamurthy, Kendall Lowrey, Motoya Ohnishi, and Wen Sun · 2020
Later among the works it cites.
Active learning for nonlinear system identification with guarantees
Horia Mania, Michael I Jordan, and Benjamin Recht · 2020
Later among the works it cites.
Learning the linear quadratic regulator from nonlinear observations
Zakaria Mhammedi, Dylan J Foster, Max Simchowitz, Dipendra Misra, Wen Sun, Akshay Krishnamurthy, Alexander Rakhlin, and John Langford · 2020
Later among the works it cites.
Sample complexity of reinforcement learning using linearly combined model ensembles
Aditya Modi, Nan Jiang, Ambuj Tewari, and Satinder Singh · 2020
Later among the works it cites.
Deep dynamics models for learning dexterous manipulation
Anusha Nagabandi, Kurt Konolige, Sergey Levine, and Vikash Kumar · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nan Jiang and Alekh Agarwal · 2018
Cited alongside, same era.
High-dimensional probability: An introduction with applications in data science , volume 47
Roman Vershynin · 2018
Cited alongside, same era.
Online control with adversarial disturbances
Naman Agarwal, Brian Bullins, Elad Hazan, Sham Kakade, and Karan Singh · 2019
Cited alongside, same era.
Provably efficient rl with rich observations via latent state decoding
Simon Du, Akshay Krishnamurthy, Nan Jiang, Alekh Agarwal, Miroslav Dudik, and John Langford · 2019
Cited alongside, same era.
Certainty equivalence is efficient for linear quadratic control
Horia Mania, Stephen Tu, and Benjamin Recht · 2019
Cited alongside, same era.
Kinematic state abstraction and provably efficient rich-observation reinforcement learning
Dipendra Misra, Mikael Henaff, Akshay Krishnamurthy, and John Langford · 2019
Cited alongside, same era.
Worst-case regret bounds for exploration via randomized value functions
Daniel Russo · 2019
Cited alongside, same era.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Later among the works it cites.
Planning to explore via self-supervised world models
Ramanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel, Danijar Hafner, and Deepak Pathak · 2020
Later among the works it cites.
Naive exploration is optimal for online lqr
Max Simchowitz and Dylan Foster · 2020
Later among the works it cites.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
Lin Yang and Mengdi Wang · 2020
Later among the works it cites.
Frequentist regret bounds for randomized least-squares value iteration
Andrea Zanette, David Brandfonbrener, Emma Brunskill, Matteo Pirotta, and Alessandro Lazaric · 2020
Later among the works it cites.
Bilinear classes: A structural framework for provable generalization in rl
Simon S Du, Sham M Kakade, Jason D Lee, Shachar Lovett, Gaurav Mahajan, Wen Sun, and Ruosong Wang · 2021
Later among the works it cites.
The statistical complexity of interactive decision making
Dylan J Foster, Sham M Kakade, Jian Qian, and Alexander Rakhlin · 2021
Later among the works it cites.
Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms
Chi Jin, Qinghua Liu, and Sobhan Miryoosefi · 2021
Later among the works it cites.
Representation learning for online and offline rl in low-rank mdps
Masatoshi Uehara, Xuezhou Zhang, and Wen Sun · 2021
Later among the works it cites.
Feel-good thompson sampling for contextual bandits and reinforcement learning
Tong Zhang · 2021
Later among the works it cites.
Provably efficient representation learning in low-rank markov decision processes
Weitong Zhang, Jiafan He, Dongruo Zhou, Amy Zhang, and Quanquan Gu · 2021
Later among the works it cites.
Nearly minimax optimal reinforcement learning for linear mixture markov decision processes
Dongruo Zhou, Quanquan Gu, and Csaba Szepesvari · 2021
Later among the works it cites.
Alekh Agarwal and Tong Zhang · 2022
Closest in time.