Fetching the paper…
Reading the bibliography…
Learning to plan for long horizons is a central challenge in episodic reinforcement learning problems.
Provably efficient RL with rich observations via latent state decoding
Simon S Du, Akshay Krishnamurthy, Nan Jiang, Alekh Agarwal, Miroslav Dudík, and John Langford · 1901
Earlier work this paper cites.
Is a good representation sufficient for sample efficient reinforcement learning?
Simon S Du, Sham M Kakade, Ruosong Wang, and Lin F Yang · 1910
Earlier work this paper cites.
An upper bound on the loss from approximate optimal-value functions
Satinder P Singh and Richard C Yee · 1994
Earlier work this paper cites.
Finite-sample convergence rates for Q-learning and indirect algorithms
Michael J Kearns and Satinder P Singh · 1999
Earlier work this paper cites.
Approximate planning in large POMDPs via reusable trajectories
Michael J Kearns, Yishay Mansour, and Andrew Y Ng · 2000
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham M Kakade · 2003
Earlier work this paper cites.
Knows what it knows: a framework for self-aware learning
Lihong Li, Michael L Littman, Thomas J Walsh, and Alexander L Strehl · 2011
Earlier work this paper cites.
Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert J Kappen · 2013
Earlier work this paper cites.
Batch mode reinforcement learning based on the synthesis of artificial trajectories
Raphael Fonteneau, Susan A Murphy, Louis Wehenkel, and Damien Ernst · 2013
Earlier work this paper cites.
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Earlier work this paper cites.
PAC reinforcement learning with rich observations
Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2016
Earlier work this paper cites.
On lower bounds for regret in reinforcement learning
Ian Osband and Benjamin Van Roy · 2016
Cited alongside, same era.
Efficient reinforcement learning in deterministic systems with value function generalization
Zheng Wen and Benjamin Van Roy · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Unifying PAC and regret: Uniform PAC bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Cited alongside, same era.
Contextual decision processes with low Bellman rank are PAC-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Cited alongside, same era.
A theoretical analysis of deep Q-learning
Jianqing Fan, Zhaoran Wang, Yuchen Xie, and Zhuoran Yang · 2019
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2019
Later among the works it cites.
Learning with good feature representations in bandits and in RL with a generative model
Tor Lattimore and Csaba Szepesvari · 2019
Later among the works it cites.
Comments on the Du-Kakade-Wang-Yang lower bounds
Benjamin Van Roy and Shi-Hai Dong · 2019
Later among the works it cites.
Reinforcement leaning in feature space: Matrix bandit, kernels, and regret bound
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ian Osband and Benjamin Van Roy · 2017
Cited alongside, same era.
On oracle-efficient PAC RL with rich observations
Christoph Dann, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E. Schapire · 2018
Cited alongside, same era.
Open problem: The dependence of sample complexity lower bounds on planning horizon
Nan Jiang and Alekh Agarwal · 2018
Cited alongside, same era.
Is Q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Cited alongside, same era.
On the optimality of sparse model-based planning for Markov decision processes
Alekh Agarwal, Sham Kakade, and Lin F Yang · 2019
Cited alongside, same era.
Policy certificates: Towards accountable reinforcement learning
Christoph Dann, Lihong Li, Wei Wei, and Emma Brunskill · 2019
Cited alongside, same era.
Provably efficient Q-learning with function approximation via distribution shift error checking oracle
Simon S Du, Yuping Luo, Ruosong Wang, and Hanrui Zhang
Cited in the paper.
Lin F Yang and Mengdi Wang · 2019
Later among the works it cites.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Andrea Zanette and Emma Brunskill · 2019
Later among the works it cites.
PAC reinforcement learning without real-world feedback
Yuren Zhong, Aniket Anand Deshmukh, and Clayton Scott · 2019
Later among the works it cites.
Simon S Du, Jason D Lee, Gaurav Mahajan, and Ruosong Wang · 2020
Closest in time.
Provably efficient exploration for RL with unsupervised learning
Fei Feng, Ruosong Wang, Wotao Yin, Simon S Du, and Lin F Yang · 2020
Closest in time.
Learning near optimal policies with low inherent Bellman error
Andrea Zanette, Alessandro Lazaric, Mykel Kochenderfer, and Emma Brunskill · 2020
Closest in time.