Task-level robot learning: Juggling a tennis ball more accurately
Eric W Aboaf, Steven M Drucker, and Christopher G Atkeson · 1989
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Integral probability metrics and their generating classes of functions
Alfred Müller · 1997
Earlier work this paper cites.
Efficient reinforcement learning in factored MDPs
Michael Kearns and Daphne Koller · 1999
Earlier work this paper cites.
Complexity of finite-horizon Markov decision process problems
Martin Mundhenk, Judy Goldsmith, Christopher Lusena, and Eric Allender · 2000
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
A note on the representational incompatibility of function approximation and factored dynamics
Eric Allender, Sanjeev Arora, Michael Kearns, Cristopher Moore, and Alexander Russell · 2003
Earlier work this paper cites.
Efficient solution algorithms for factored MDPs
Carlos Guestrin, Daphne Koller, Ronald Parr, and Shobha Venkataraman · 2003
Earlier work this paper cites.
Exploration in metric state spaces
Sham Kakade, Michael J Kearns, and John Langford · 2003
Earlier work this paper cites.
PAC model-free reinforcement learning
Alexander L Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L Littman · 2006
Earlier work this paper cites.
Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
The adaptive k-meteorologists problem and its application to structure learning and feature selection in reinforcement learning
Carlos Diuk, Lihong Li, and Bethany R Leffler · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.