Neuro-dynamic programming
D. P. Bertsekas and J. N. Tsitsiklis · 1996
Earlier work this paper cites.
Agnostic q-learning with function approximation in deterministic systems: Tight bounds on approximation error and sample complexity
Original
S. S. Du, J. D. Lee, G. Mahajan, and R. Wang · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
M. Kearns and S. Singh · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
S. M. Kakade · 2003
Earlier work this paper cites.
Error bounds for approximate policy iteration
R. Munos · 2003
Earlier work this paper cites.
Finite time bounds for sampling based fitted value iteration
C. Szepesvári and R. Munos · 2005
Earlier work this paper cites.
Sampling algorithms for l 2 regression and applications
P. Drineas, M. W. Mahoney, and S. Muthukrishnan · 2006
Earlier work this paper cites.
PAC model-free reinforcement learning
A. L. Strehl, L. Li, E. Wiewiora, J. Langford, and M. L. Littman · 2006
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
A. Antos, C. Szepesvári, and R. Munos · 2008
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
V. Dani, T. P. Hayes, and S. M. Kakade · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
R. Munos and C. Szepesvári · 2008
Earlier work this paper cites.
Reinforcement learning in finite MDPs: PAC analysis
A. L. Strehl, L. Li, and M. L. Littman · 2009
Earlier work this paper cites.
Parametric bandits: The generalized linear case
S. Filippi, O. Cappe, A. Garivier, and C. Szepesvári · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Earlier work this paper cites.
Universal ε \varepsilon -approximators for integrals
M. Langberg and L. J. Schulman · 2010
Earlier work this paper cites.
Model-based reinforcement learning with nearly tight exploration complexity bounds
I. Szita and C. Szepesvári · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Y. Abbasi-Yadkori, D. Pál, and C. Szepesvári · 2011
Earlier work this paper cites.
A unified framework for approximating and clustering data
D. Feldman and M. Langberg · 2011
Earlier work this paper cites.
Graph sparsification by effective resistances
D. A. Spielman and N. Srivastava · 2011
Earlier work this paper cites.
Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
M. G. Azar, R. Munos, and H. J. Kappen · 2013
Earlier work this paper cites.
Turning big data into tiny data: Constant-size coresets for k-means, pca and projective clustering
D. Feldman, M. Schmidt, and C. Sohler · 2013
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
J. Kober, J. A. Bagnell, and J. Peters · 2013
Earlier work this paper cites.