Fetching the paper…
Reading the bibliography…
We consider approximate dynamic programming for the infinite-horizon stationary $\gamma$-discounted optimal control problem formalized by Markov Decision Processes.
Modified policy iteration algorithms for discounted Markov decision problems
Puterman, M. and Shin, M. (1978) · 1978
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Puterman, M. (1994) · 1994
Earlier work this paper cites.
An upper bound on the loss from approximate optimal-value functions
Singh, S. and Yee, R. (1994) · 1994
Earlier work this paper cites.
Neuro-dynamic programming
Bertsekas, D. and Tsitsiklis, J. (1996) · 1996
Earlier work this paper cites.
On the Sample Complexity of Reinforcement Learning
Kakade, S. (2003) · 2003
Earlier work this paper cites.
Error bounds for approximate policy iteration
Munos, R. (2003) · 2003
Cited alongside, same era.
Value-iteration based fitted policy iteration: learning with a single trajectory
Antos, A., Szepesvarf, C., and Munos, R. (2007b) · 2007
Cited alongside, same era.
Performance bounds in L p {L}_{p} -norm for approximate value iteration
Munos, R. (2007) · 2007
Cited alongside, same era.
Error propagation for approximate policy and value iteration (extended version)
Farahmand, A., Munos, R., and Szepesvári, C. (2010) · 2010
Cited alongside, same era.
Fitted Q-iteration in continuous action-space MDPs
Antos, A., Munos, R., Szepesvari, C., et al
Cited in the paper.
Approximate modified policy iteration
Scherrer, B., Ghavamzadeh, M., Gabillon, V., and Geist, M. (2012a)
Cited in the paper.
Approximate modified policy iteration
Scherrer, B., Gabillon, V., Ghavamzadeh, M., and Geist, M. (2012b)
Cited in the paper.
Performance bound for approximate optimistic policy iteration
Scherrer, B. and Thiery, C. (2010) · 2010
Later among the works it cites.
Q-learning and enhanced policy iteration in discounted dynamic programming
Bertsekas, D. and Yu, H. (2012) · 2012
Later among the works it cites.
(Approximate) iterated successive approximations algorithm for sequential decision processes
Canbolat, P. and Rothblum, U. (2012) · 2012
Later among the works it cites.
On the use of non-stationary policies for stationary infinite-horizon markov decision processes
Scherrer, B. and Lesner, B. (2012) · 2012
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…