Fetching the paper…
Reading the bibliography…
We study query and computationally efficient planning algorithms with linear function approximation and a simulator.
An upper bound on the loss from approximate optimal-value functions
Satinder P Singh and Richard C Yee · 1994
Earlier work this paper cites.
Matrix algebra from a statistician’s perspective, 1998
David A Harville · 1998
Earlier work this paper cites.
Finite-sample convergence rates for Q-learning and indirect algorithms
Michael Kearns and Satinder Singh · 1999
Earlier work this paper cites.
A sparse sampling algorithm for near-optimal planning in large Markov decision processes
Michael Kearns, Yishay Mansour, and Andrew Y Ng · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade · 2003
Earlier work this paper cites.
Error bounds for approximate policy iteration
Rémi Munos · 2003
Earlier work this paper cites.
On the generalization ability of on-line learning algorithms
Nicolo Cesa-Bianchi, Alex Conconi, and Claudio Gentile · 2004
Earlier work this paper cites.
Prediction, learning, and games
Nicolo Cesa-Bianchi and Gábor Lugosi · 2006
Earlier work this paper cites.
PC-PG: Policy cover directed exploration for provable policy gradient learning
Alekh Agarwal, Mikael Henaff, Sham Kakade, and Wen Sun · 2007
Earlier work this paper cites.
Online linear optimization and adaptive routing
Baruch Awerbuch and Robert Kleinberg · 2008
Earlier work this paper cites.
Online Markov decision processes
Eyal Even-Dar, Sham M Kakade, and Yishay Mansour · 2009
Earlier work this paper cites.
Error propagation for approximate policy and value iteration
Amir Massoud Farahmand, Rémi Munos, and Csaba Szepesvári · 2010
Earlier work this paper cites.
Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert J Kappen · 2013
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Daniel Russo and Benjamin Van Roy · 2013
Earlier work this paper cites.
Efficient exploration and value function generalization in deterministic systems
Zheng Wen and Benjamin Van Roy · 2013
Earlier work this paper cites.
From bandits to Monte-Carlo tree search: The optimistic principle applied to optimization and planning
Rémi Munos · 2014
Cited alongside, same era.
Minimum-volume ellipsoids: Theory and algorithms
Michael J Todd · 2016
Cited alongside, same era.
Contextual decision processes with low Bellman rank are PAC-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Cited alongside, same era.
Near-optimal time and sample complexities for solving Markov decision processes with a generative model
Aaron Sidford, Mengdi Wang, Xian Wu, Lin Yang, and Yinyu Ye · 2018
Cited alongside, same era.
Politex: Regret bounds for policy iteration using expert prediction
Yasin Abbasi-Yadkori, Peter Bartlett, Kush Bhatia, Nevena Lazic, Csaba Szepesvari, and Gellért Weisz · 2019
Cited alongside, same era.
Learning with good feature representations in bandits and in RL with a generative model
Tor Lattimore, Csaba Szepesvari, and Gellert Weisz · 2020
Later among the works it cites.
Breaking the sample size barrier in model-based reinforcement learning with a generative model
Gen Li, Yuting Wei, Yuejie Chi, Yuantao Gu, and Yuxin Chen · 2020
Later among the works it cites.
Efficient planning in large MDPs with weak linear function approximation
Roshan Shariff and Csaba Szepesvári · 2020
Later among the works it cites.
Reinforcement leaning in feature space: Matrix bandit, kernels, and regret bound
Lin Yang and Mengdi Wang · 2020
Later among the works it cites.
Learning near optimal policies with low inherent Bellman error
Andrea Zanette, Alessandro Lazaric, Mykel Kochenderfer, and Emma Brunskill · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Simon S Du, Yuping Luo, Ruosong Wang, and Hanrui Zhang · 2019
Cited alongside, same era.
Go-explore: a new approach for hard-exploration problems
Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth O Stanley, and Jeff Clune · 2019
Cited alongside, same era.
Comments on the Du-Kakade-Wang-Yang lower bounds
Benjamin Van Roy and Shi Dong · 2019
Cited alongside, same era.
Optimism in reinforcement learning with generalized linear function approximation
Yining Wang, Ruosong Wang, Simon S Du, and Akshay Krishnamurthy · 2019
Cited alongside, same era.
Sample-optimal parametric Q-learning using linearly additive features
Lin Yang and Mengdi Wang · 2019
Cited alongside, same era.
Limiting extrapolation in linear approximate value iteration
Andrea Zanette, Alessandro Lazaric, Mykel J Kochenderfer, and Emma Brunskill · 2019
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvari, Mengdi Wang, and Lin Yang · 2020
Cited alongside, same era.
Dongruo Zhou, Jiafan He, and Quanquan Gu · 2020
Later among the works it cites.
Bilinear classes: A structural framework for provable generalization in RL
Simon S Du, Sham M Kakade, Jason D Lee, Shachar Lovett, Gaurav Mahajan, Wen Sun, and Ruosong Wang · 2021
Closest in time.
Adaptive approximate policy iteration
Botao Hao, Nevena Lazic, Yasin Abbasi-Yadkori, Pooria Joulani, and Csaba Szepesvári · 2021
Closest in time.
Bellman eluder dimension: New rich classes of RL problems, and sample-efficient algorithms
Chi Jin, Qinghua Liu, and Sobhan Miryoosefi · 2021
Closest in time.
Improved regret bound and experience replay in regularized policy iteration
Nevena Lazic, Dong Yin, Yasin Abbasi-Yadkori, and Csaba Szepesvari · 2021
Closest in time.
Gen Li, Yuxin Chen, Yuejie Chi, Yuantao Gu, and Yuting Wei · 2021
Closest in time.
RL Theory lecture notes: POLITEX
Csaba Szepesvári · 2021
Closest in time.
An exponential lower bound for linearly-realizable MDPs with constant suboptimality gap
Yuanhao Wang, Ruosong Wang, and Sham M Kakade · 2021
Closest in time.
Learning infinite-horizon average-reward MDPs with linear function approximation
Chen-Yu Wei, Mehdi Jafarnia Jahromi, Haipeng Luo, and Rahul Jain · 2021
Closest in time.
Cautiously optimistic policy optimization and exploration with linear function approximation
Andrea Zanette, Ching-An Cheng, and Alekh Agarwal · 2021
Closest in time.