Fetching the paper…
Reading the bibliography…
We consider approximate dynamic programming in $\gamma$-discounted Markov decision processes and apply it to approximate planning with linear value-function approximation.
Inverting modified matrices
A Woodbury Max · 1950
Earlier work this paper cites.
Principles of mathematical analysis , volume 3
Walter Rudin et al · 1976
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Neuro-Dynamic Programming
Dimitri P. Bertsekas and John N. Tsitsiklis · 1996
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Error bounds for approximate policy iteration
Remi Munos · 2003
Earlier work this paper cites.
Error bounds for approximate value iteration
Remi Munos · 2005
Earlier work this paper cites.
Dynamic Programming and Optimal Control: Approximate dynamic programming , volume II
Dimitri P. Bertsekas · 2012
Earlier work this paper cites.
On the use of non-stationary policies for stationary infinite-horizon Markov decision processes
Bruno Scherrer and Boris Lesner · 2012
Cited alongside, same era.
Approximate policy iteration schemes: a comparison
Bruno Scherrer · 2014
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Is a good representation sufficient for sample efficient reinforcement learning?
Simon S Du, Sham M Kakade, Ruosong Wang, and Lin F Yang · 2019
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Cited alongside, same era.
Bandit Algorithms
T. Lattimore and Cs. Szepesvári · 2020
Cited alongside, same era.
Learning near optimal policies with low inherent Bellman error
Andrea Zanette, Alessandro Lazaric, Mykel Kochenderfer, and Emma Brunskill · 2020
Later among the works it cites.
Provably efficient reinforcement learning for discounted MDPs with feature mapping
Dongruo Zhou, Jiafan He, and Quanquan Gu · 2020
Later among the works it cites.
An exponential lower bound for linearly realizable MDP with constant suboptimality gap
Yuanhao Wang, Ruosong Wang, and Sham Kakade · 2021
Later among the works it cites.
Exponential lower bounds for planning in MDPs with linearly-realizable optimal action-value functions
Gellért Weisz, Philip Amortila, and Csaba Szepesvári · 2021
Later among the works it cites.
Reward-free RL is no harder than reward-aware RL in linear Markov decision processes
Andrew Wagenmaker, Yifang Chen, Max Simchowitz, Simon S Du, and Kevin Jamieson · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning with good feature representations in bandits and in RL with a generative model
Tor Lattimore, Csaba Szepesvári, and Gellért Weisz · 2020
Cited alongside, same era.
Approximation benefits of policy gradient methods with aggregated states
Daniel Russo · 2020
Cited alongside, same era.
Closest in time.
The curse of passive data collection in batch reinforcement learning
Chenjun Xiao, Ilbin Lee, Bo Dai, Dale Schuurmans, and Csaba Szepesvari · 2022
Closest in time.
Efficient local planning with linear function approximation
Dong Yin, Botao Hao, Yasin Abbasi-Yadkori, Nevena Lazić, and Csaba Szepesvári · 2022
Closest in time.