Efficient reinforcement learning in factored mdps
Michael Kearns and Daphne Koller · 1999
Earlier work this paper cites.
Policy iteration for factored mdps
Daphne Koller and Ronald Parr · 2000
Earlier work this paper cites.
Empirical Processes in M-Estimation
S van de Geer · 2000
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Earlier work this paper cites.
Robust dynamic programming
Garud N Iyengar · 2005
Earlier work this paper cites.
Error bounds for approximate value iteration
Rémi Munos · 2005
Earlier work this paper cites.
Robust control of markov decision processes with uncertain transition matrices
Arnab Nilim and Laurent El Ghaoui · 2005
Earlier work this paper cites.
Gaussian Processes for Machine Learning (Adaptive Computation and Machine Learning)
Carl Edward Rasmussen and Christopher K. I. Williams · 2005
Earlier work this paper cites.
Learning bounds for kernel regression using effective data dimensionality
Tong Zhang · 2005
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Niranjan Srinivas, Andreas Krause, Sham Kakade, and Matthias Seeger · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Marc Deisenroth and Carl E Rasmussen · 2011
Earlier work this paper cites.
Is pessimism provably efficient for offline rl?
Original
Ying Jin, Zhuoran Yang, and Zhaoran Wang · 2012
Earlier work this paper cites.
Agnostic system identification for model-based reinforcement learning
Stéphane Ross and J Andrew Bagnell · 2012
Earlier work this paper cites.
Finite-time analysis of kernelised contextual bandits
Michal Valko, Nathan Korda, Rémi Munos, Ilias Flaounas, and Nello Cristianini · 2013
Earlier work this paper cites.
Learning to optimize via posterior sampling
Daniel Russo and Benjamin Van Roy · 2014
Earlier work this paper cites.
Scaling up robust mdps using function approximation
Aviv Tamar, Shie Mannor, and Huan Xu · 2014
Earlier work this paper cites.
On the equivalence between kernel quadrature rules and random feature expansions
Francis Bach · 2017
Earlier work this paper cites.
The total variation distance between high-dimensional gaussians
Original
Luc Devroye, Abbas Mehrabian, and Tommy Reddad · 2018
Earlier work this paper cites.
Reinforcement learning: Theory and algorithms
Alekh Agarwal, Nan Jiang, Sham M Kakade, and Wen Sun · 2019
Earlier work this paper cites.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Earlier work this paper cites.
Top-k off-policy correction for a reinforce recommender system
Minmin Chen, Alex Beutel, Paul Covington, Sagar Jain, Francois Belletti, and Ed Chi · 2019
Earlier work this paper cites.