Fetching the paper…
Reading the bibliography…
Consider a Markov decision process (MDP) that admits a set of state-action features, which can linearly express the process's probabilistic transition model.
An upper bound on the loss from approximate optimal-value functions
Satinder P Singh and Richard C Yee · 1994
Earlier work this paper cites.
Reinforcement learning with soft state aggregation
Satinder P Singh, Tommi Jaakkola, and Michael I Jordan · 1995
Earlier work this paper cites.
Feature-based methods for large scale dynamic programming
John N Tsitsiklis and Benjamin Van Roy · 1996
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
John N Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
Finite-sample convergence rates for q-learning and indirect algorithms
Michael J Kearns and Satinder P Singh · 1999
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade · 2003
Earlier work this paper cites.
Least-squares policy iteration
Michail G Lagoudakis and Ronald Parr · 2003
Earlier work this paper cites.
Least squares policy evaluation algorithms with linear function approximation
A Nedić and Dimitri P Bertsekas · 2003
Earlier work this paper cites.
When does non-negative matrix factorization give a correct decomposition into parts?
David Donoho and Victoria Stodden · 2004
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Dimitri P Bertsekas · 2005
Cited alongside, same era.
Action elimination and stopping conditions for the multi-armed bandit and reinforcement learning problems
Eyal Even-Dar, Shie Mannor, and Yishay Mansour · 2006
Cited alongside, same era.
An analysis of reinforcement learning with function approximation
Francisco S Melo, Sean P Meyn, and M Isabel Ribeiro · 2008
Cited alongside, same era.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Cited alongside, same era.
An analysis of linear models, linear value-function approximation, and feature selection for reinforcement learning
Ronald Parr, Lihong Li, Gavin Taylor, Christopher Painter-Wakefield, and Michael L Littman · 2008
Cited alongside, same era.
A convergent o ( n ) o(n) temporal-difference algorithm for off-policy learning with linear function approximation
Finite-sample analysis of least-squares policy iteration
Alessandro Lazaric, Mohammad Ghavamzadeh, and Rémi Munos · 2012
Later among the works it cites.
Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert J Kappen · 2013
Later among the works it cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Later among the works it cites.
On the rate of convergence and error bounds for LSTD( λ \lambda )
Manel Tagorti and Bruno Scherrer · 2015
Later among the works it cites.
Reinforcement learning in rich-observation mdps using spectral methods
Kamyar Azizzadenesheli, Alessandro Lazaric, and Animashree Anandkumar · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Richard S Sutton, Hamid R Maei, and Csaba Szepesvári · 2009
Cited alongside, same era.
Error propagation for approximate policy and value iteration
Amir-massoud Farahmand, Csaba Szepesvári, and Rémi Munos · 2010
Cited alongside, same era.
Toward off-policy learning control with function approximation
Hamid Reza Maei, Csaba Szepesvári, Shalabh Bhatnagar, and Richard S Sutton · 2010
Cited alongside, same era.
Learning topic models–going beyond svd
Sanjeev Arora, Rong Ge, and Ankur Moitra · 2012
Cited alongside, same era.
Fitted Q-iteration in continuous action-space mdps
András Antos, Csaba Szepesvári, and Rémi Munos
Cited in the paper.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos
Cited in the paper.
Reinforcement learning with a near optimal rate of convergence
Mohammad Gheshlaghi Azar, Rémi Munos, Mohammad Ghavamzadeh, and Hilbert Kappen
Cited in the paper.
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Later among the works it cites.
Scalable bilinear pi learning using state and action features
Yichen Chen, Lihong Li, and Mengdi Wang · 2018
Later among the works it cites.
State aggregation learning from markov transition data
Yaqi Duan, Zheng Tracy Ke, and Mengdi Wang · 2018
Later among the works it cites.