Provably efficient RL with rich observations via latent state decoding
Original
Simon S Du, Akshay Krishnamurthy, Nan Jiang, Alekh Agarwal, Miroslav Dudík, and John Langford · 1901
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
The central limit theorem for weighted empirical processes indexed by sets
Kenneth S Alexander · 1987
Earlier work this paper cites.
On learning sets and functions
Balas K Natarajan · 1989
Earlier work this paper cites.
Optimal adaptive policies for sequential allocation problems
Apostolos N Burnetas and Michael N Katehakis · 1996
Earlier work this paper cites.
Associative reinforcement learning using linear probabilistic concepts
Naoki Abe and Philip M Long · 1999
Earlier work this paper cites.
Beating the hold-out: Bounds for k-fold and progressive cross-validation
Avrim Blum, Adam Kalai, and John Langford · 1999
Earlier work this paper cites.
Smooth discrimination analysis
Enno Mammen and Alexandre B Tsybakov · 1999
Earlier work this paper cites.
Agnostic Q-learning with function approximation in deterministic systems: Tight bounds on approximation error and sample complexity
Original
Simon S Du, Jason D Lee, Gaurav Mahajan, and Ruosong Wang · 2002
Earlier work this paper cites.
Reinforcement learning with immediate rewards and linear hypotheses
Naoki Abe, Alan W Biermann, and Philip M Long · 2003
Earlier work this paper cites.
Optimal aggregation of classifiers in statistical learning
Alexander B Tsybakov · 2004
Earlier work this paper cites.
Concentration inequalities and asymptotic results for ratio type empirical processes
Evarist Giné and Vladimir Koltchinskii · 2006
Earlier work this paper cites.
Fast learning rates for plug-in classifiers
Jean-Yves Audibert, Alexandre B Tsybakov, et al · 2007
Earlier work this paper cites.
A bound on the label complexity of agnostic active learning
Steve Hanneke · 2007
Earlier work this paper cites.
Progressive mixture rules are deviation suboptimal
Jean-Yves Audibert · 2008
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas P Hayes, and Sham M Kakade · 2008
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
John Langford and Tong Zhang · 2008
Earlier work this paper cites.
Optimistic linear programming gives logarithmic regret for irreducible MDPs
Ambuj Tewari and Peter L Bartlett · 2008
Earlier work this paper cites.
Introduction to Nonparametric Estimation
Alexandre B Tsybakov · 2008
Earlier work this paper cites.
Active learning for smooth problems
Eric Friedman · 2009
Earlier work this paper cites.
The true sample complexity of active learning
Maria-Florina Balcan, Steve Hanneke, and Jennifer Wortman Vaughan · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Earlier work this paper cites.
Nonparametric bandits with covariates
Philippe Rigollet and Assaf Zeevi · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.