Frequentist regret bounds for randomized least-squares value iteration
Andrea Zanette, David Brandfonbrener, Emma Brunskill, Matteo Pirotta, and Alessandro Lazaric · 1964
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Stochastic processes with sample paths in reproducing kernel Hilbert spaces
M. N. Lukic and J. H. Beder · 2001
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
Bernhard Schölkopf and Alexander J. Smola · 2002
Earlier work this paper cites.
Learning near optimal policies with low inherent bellman error
Original
Andrea Zanette, Alessandro Lazaric, Mykel Kochenderfer, and Emma Brunskill · 2003
Earlier work this paper cites.
On the nyström method for approximating a gram matrix for improved kernel-based learning
Petros Drineas and Michael W Mahoney · 2005
Earlier work this paper cites.
Gaussian processes for machine learning
C. E. Rasmussen and C. K. I. Williams · 2006
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
Kernel measures of conditional dependence
Kenji Fukumizu, Arthur Gretton, Xiaohai Sun, and Bernhard Schölkopf · 2008
Earlier work this paper cites.
Support vector machines
Ingo Steinwart and Andreas Christmann · 2008
Earlier work this paper cites.
Kernel dimension reduction in regression
Kenji Fukumizu, Francis R Bach, Michael I Jordan, et al · 2009
Earlier work this paper cites.
Hilbert space embeddings of conditional distributions with applications to dynamical systems
Le Song, Jonathan Huang, Alex Smola, and Kenji Fukumizu · 2009
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Original
Niranjan Srinivas, Andreas Krause, Sham M Kakade, and Matthias Seeger · 2009
Earlier work this paper cites.
Reinforcement learning in finite mdps: PAC
Alexander L Strehl, Lihong Li, and Michael L Littman · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Regret bounds for the adaptive control of linear quadratic systems
Yasin Abbasi-Yadkori and Csaba Szepesvári · 2011
Earlier work this paper cites.