Fetching the paper…
Reading the bibliography…
We investigate projection methods, for evaluating a linear approximation of the value function of a policy in a Markov Decision Process context.
Tight performance bounds on greedy policies based on imperfect value functions
Williams, R. J. and Baird, L. C · 1993
Earlier work this paper cites.
Stable function approximation in dynamic programming
Gordon, G · 1995
Earlier work this paper cites.
Neurodynamic Programming
Bertsekas, D.P. and Tsitsiklis, J.N · 1996
Earlier work this paper cites.
Minkowski Geometry
Thompson, A.C · 1996
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
Tsitsiklis, J.N. and Van Roy, B · 1997
Earlier work this paper cites.
Max-norm projections for factored mdps
Guestrin, C., Koller, D., and Parr, R · 2001
Earlier work this paper cites.
Technical update: Least-squares temporal difference learning
Boyan, J. A · 2002
Cited alongside, same era.
Optimality of reinforcement learning algorithms with linear function approximation
Schoknecht, R · 2002
Cited alongside, same era.
Least-squares policy iteration
Lagoudakis, M. G. and Parr, R · 2003
Cited alongside, same era.
Error bounds for approximate policy iteration
Munos, R · 2003
Cited alongside, same era.
Iterative Methods for Sparse Linear Systems, 2nd edition
Saad, Y · 2003
Cited alongside, same era.
The many proofs of an identity on the norm of oblique projections
Szyld, D.B · 2006
Cited alongside, same era.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
Antos, A., Szepesvári, C., and Munos, R · 2008
Later among the works it cites.
Regularized policy iteration
Farahmand, A.M., Ghavamzadeh, M., Szepesvári, C., and Mannor, S · 2008
Later among the works it cites.
Finite-time bounds for fitted value iteration
Munos, R. and Szepesvári, C · 2008
Later among the works it cites.
New error bounds for approximations from projected linear equations
Yu, H. and Bertsekas, D.P · 2008
Later among the works it cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Sutton, R. S., Maei, H. R., Precup, D., Bhatnagar, S., Silver, D., Szepesvári, C., and Wiewiora, E · 2009
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…