Fetching the paper…
Reading the bibliography…
We consider off-policy temporal-difference (TD) learning methods for policy evaluation in Markov decision processes with finite spaces and discounted reward criteria, and we present a collection of convergence results for several gradient-based TD algorithms with linear function approximation.
Real and Complex Analysis
Rudin, W. (1966) · 1966
Earlier work this paper cites.
Convex Analysis
Rockafellar, R. T. (1970) · 1970
Earlier work this paper cites.
Stochastic Approximation Methods for Constrained and Unconstrained Systems
Kushner, H. J. and Clark, D. S. (1978) · 1978
Earlier work this paper cites.
Problem Complexity and Method Efficiency in Optimization
Nemirovsky, A. S. and Yudin, D. B. (1983) · 1983
Earlier work this paper cites.
General Irreducible Markov Chains and Non-Negative Operators
Nummelin, E. (1984) · 1984
Earlier work this paper cites.
Differential Inclusions: Set-Valued Maps and Viability Theory
Aubin, J.-P. and Cellina, A. (1985) · 1985
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S. (1988) · 1988
Earlier work this paper cites.
Ergodic theorems for discrete time stochastic systems using a stochastic Lyapunov function
Meyn, S. (1989) · 1989
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Polyak, B. T. and Juditsky, A. B. (1992) · 1992
Earlier work this paper cites.
The General Topology of Dynamical Systems
Akin, E. (1993) · 1993
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L. (1994) · 1994
Earlier work this paper cites.
TD models: Modeling the world at a mixture of time scales
Sutton, R. S. (1995) · 1995
Earlier work this paper cites.
Neuro-Dynamic Programming
Bertsekas, D. P. and Tsitsiklis, J. N. (1996) · 1996
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
Tsitsiklis, J. N. and Van Roy, B. (1997) · 1997
Earlier work this paper cites.
Variational Analysis
Rockafellar, R. T. and Wets, R. J.-B. (1998) · 1998
Cited alongside, same era.
Reinforcement Learning
Sutton, R. S. and Barto, A. G. (1998) · 1998
Cited alongside, same era.
Eligibility traces for off-policy policy evaluation
Precup, D., Sutton, R. S., and Singh, S. (2000) · 2000
Cited alongside, same era.
Real Analysis and Probability
Dudley, R. M. (2002) · 2002
Cited alongside, same era.
Stochastic Approximation and Recursive Algorithms and Applications
Kushner, H. J. and Yin, G. G. (2003) · 2003
Cited alongside, same era.
Stochastic Approximation: A Dynamic Viewpoint
Borkar, V. S. (2008) · 2008
Cited alongside, same era.
A convergent O ( n ) O(n) algorithm for off-policy temporal-difference learning with linear function approximation
Gradient Temporal-Difference Learning Algorithms
Maei, H. R. (2011) · 2011
Later among the works it cites.
Sparse Q-learning with mirror descent
Mahadevan, S. and Liu, B. (2012) · 2012
Later among the works it cites.
Least squares temporal difference methods: An analysis under general conditions
Yu, H. (2012) · 2012
Later among the works it cites.
Weighted Bellman equations and their applications in approximate dynamic programming
Yu, H. and Bertsekas, D. P. (2012) · 2012
Later among the works it cites.
Proximal reinforcement learning: A new theory of sequential decision making in primal-dual spaces
Mahadevan, S., Liu, B., Thomas, P., Dabney, W., Giguere, S., Jacek, N., Gemp, I., and Liu, J. (2014) · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sutton, R. S., Szepesvári, C., and Maei, H. R. (2008) · 2008
Cited alongside, same era.
Markov Chains and Stochastic Stability
Meyn, S. and Tweedie, R. L. (2009) · 2009
Cited alongside, same era.
The grand challenge of predictive empirical abstract knowledge
Sutton, R. S. (2009) · 2009
Cited alongside, same era.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Sutton, R. S., Maei, H. R., Precup, D., Bhatnagar, S., Silver, D., Szepesvári, C., and Wiewiora, E. (2009) · 2009
Cited alongside, same era.
GQ( λ \lambda ): A general gradient algorithm for temporal-difference prediction learning with eligibility traces
Maei, H. R. and Sutton, R. S. (2010) · 2010
Cited alongside, same era.
Should one compute the temporal difference fix point or minimize the Bellman residual? The unified oblique projection view
Scherrer, B. (2010) · 2010
Cited alongside, same era.
Karmakar, P. and Bhatnagar, S. (2015) · 2015
Later among the works it cites.
Finite-sample analysis of proximal gradient TD algorithms
Liu, B., Liu, J., Ghavamzadeh, M., Mahadevan, S., and Petrik, M. (2015) · 2015
Later among the works it cites.
On convergence of emphatic temporal-difference learning
Yu, H. (2015) · 2015
Later among the works it cites.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M. G. (2016) · 2016
Later among the works it cites.
An emphatic approach to the problem of off-policy temporal-difference learning
Sutton, R. S., Mahmood, A. R., and White, M. (2016) · 2016
Later among the works it cites.
Weak convergence properties of constrained emphatic temporal-difference learning with constant and slowly diminishing stepsize
Yu, H. (2016) · 2016
Later among the works it cites.
Multi-step off-policy learning without importance-sampling ratios
Mahmood, A. R., Yu, H., and Sutton, R. S. (2017) · 2017
Closest in time.
On generalized Bellman equations and temporal-difference learning
Yu, H., Mahmood, A. R., and Sutton, R. S. (2017) · 2017
Closest in time.