Fetching the paper…
Reading the bibliography…
Policy evaluation is a crucial step in many reinforcement-learning procedures, which estimates a value function that predicts states' long-term value under a given policy.
Convex Analysis
Rockafellar, R. Tyrrell · 1970
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, Long-Ji · 1992
Earlier work this paper cites.
Neuro-dynamic programming: An overview
Bertsekas, Dimitri P and Tsitsiklis, John N · 1995
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Bradtke, Steven J and Barto, Andrew G · 1996
Earlier work this paper cites.
Reinforcement learning with replacing eligibility traces
Singh, Satinder P. and Sutton, Richard S · 1996
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
Tsitsiklis, John N. and Van Roy, Benjamin · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, Richard S and Barto, Andrew G · 1998
Earlier work this paper cites.
“Bias-variance” error bounds for temporal difference updates
Kearns, Michael J. and Singh, Satinder P · 2000
Earlier work this paper cites.
Off-policy temporal-difference learning with funtion approximation
Precup, Doina, Sutton, Richard S., and Dasgupta, Sanjoy · 2001
Earlier work this paper cites.
Technical update: Least-squares temporal difference learning
Boyan, Justin A · 2002
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, Michail G and Parr, Ronald · 2003
Earlier work this paper cites.
Least squares policy evaluation algorithms with linear function approximation
Nedić, A. and Bertsekas, Dimitri P · 2003
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, Martin L · 2005
Cited alongside, same era.
Indefinite linear algebra and applications
Gohberg, Israel, Lancaster, Peter, and Rodman, Leiba · 2006
Cited alongside, same era.
iLSTD: Eligibility traces and convergence analysis
Geramifard, Alborz, Bowling, Michael H., Zinkevich, Martin, and Sutton, Richard S · 2007
Cited alongside, same era.
On nonsymmetric saddle point matrices that allow conjugate gradient iterations
Liesen, Jörg and Parlett, Beresford N · 2008
Cited alongside, same era.
A condition for the nonsymmetric saddle point matrix being diagonalizable and having real and positive eigenvalues
Shen, Shu-Qian, Huang, Ting-Zhu, and Cheng, Guang-Hui · 2008
Cited alongside, same era.
Convergent temporal-difference learning with arbitrary smooth function approximation
Policy evaluation with temporal differences: a survey and comparison
Dann, Christoph, Neumann, Gerhard, and Peters, Jan · 2014
Later among the works it cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Defazio, Aaron, Bach, Francis, and Lacoste-Julien, Simon · 2014
Later among the works it cites.
Fast LSTD using stochastic approximation: Finite time analysis and application to traffic control
Prashanth, LA, Korda, Nathaniel, and Munos, Rémi · 2014
Later among the works it cites.
On TD(0) with function approximation: Concentration bounds and a centered variant with exponential convergence
Korda, Nathaniel and Prashanth, L.A · 2015
Later among the works it cites.
A universal catalyst for first-order optimization
Lin, Hongzhou, Mairal, Julien, and Harchaoui, Zaid · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bhatnagar, Shalabh, Precup, Doina, Silver, David, Sutton, Richard S, Maei, Hamid R, and Szepesvári, Csaba · 2009
Cited alongside, same era.
Finite-sample analysis of LSTD
Lazaric, Alessandro, Ghavamzadeh, Mohammad, and Munos, Rémi · 2010
Cited alongside, same era.
A first-order primal-dual algorithm for convex problems with applications to imaging
Chambolle, Antonin and Pock, Thomas · 2011
Cited alongside, same era.
Batch reinforcement learning
Lange, Sascha, Gabel, Thomas, and Riedmiller, Martin · 2011
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, Rie and Zhang, Tong · 2013
Cited alongside, same era.
All of Statistics: A Concise Course in Statistical Inference
Wasserman, Larry · 2013
Cited alongside, same era.
A convergent o ( n ) o(n) temporal-difference algorithm for off-policy learning with linear function approximation
Sutton, Richard S, Maei, Hamid R, and Szepesvári, Csaba
Cited in the paper.
Liu, Bo, Liu, Ji, Ghavamzadeh, Mohammad, Mahadevan, Sridhar, and Petrik, Marek · 2015
Later among the works it cites.
Distributed policy evaluation under multiple behavior strategies
Valcarcel Macua, Sergio, Chen, Jianshu, Zazo, Santiago, and Sayed, Ali H · 2015
Later among the works it cites.
Stochastic variance reduction methods for saddle-point problems
Balamurugan, P and Bach, Francis · 2016
Later among the works it cites.
Learning from conditional distributions via dual embeddings
Dai, Bo, He, Niao, Pan, Yunpeng, Boots, Byron, and Song, Le · 2016
Later among the works it cites.
Accelerating stochastic composition optimization
Wang, Mengdi, Liu, Ji, and Fang, Ethan · 2016
Later among the works it cites.
Finite-sum composition optimization via variance reduced gradient descent
Lian, Xiangru, Wang, Mengdi, and Liu, Ji · 2017
Closest in time.