Fetching the paper…
Reading the bibliography…
We study the reward-free reinforcement learning framework, which is particularly suitable for batch reinforcement learning and scenarios where one needs policies for multiple reward functions.
Constrained Markov decision processes , volume 7
Eitan Altman · 1999
Earlier work this paper cites.
Finite-sample convergence rates for Q-learning and indirect algorithms
Michael J Kearns and Satinder P Singh · 1999
Earlier work this paper cites.
R-max - a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I. Brafman and Moshe Tennenholtz · 2003
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham M Kakade · 2003
Earlier work this paper cites.
Almost optimal model-free reinforcement learning via reference-advantage decomposition
Zihan Zhang, Yuan Zhou, and Xiangyang Ji · 2004
Earlier work this paper cites.
Is long horizon reinforcement learning more difficult than short horizon reinforcement learning?
Ruosong Wang, Simon S Du, Lin F Yang, and Sham M Kakade · 2005
Earlier work this paper cites.
On reward-free reinforcement learning with linear function approximation
Ruosong Wang, Simon S Du, Lin F Yang, and Ruslan Salakhutdinov · 2006
Earlier work this paper cites.
Task-agnostic exploration in reinforcement learning
Xuezhou Zhang, Adish Singla, et al · 2006
Earlier work this paper cites.
Model-free reinforcement learning: from clipped pseudo-regret to sample complexity
Zihan Zhang, Yuan Zhou, and Xiangyang Ji · 2006
Earlier work this paper cites.
Empirical Bernstein bounds and sample variance penalization
Andreas Maurer and Massimiliano Pontil · 2009
Earlier work this paper cites.
Zihan Zhang, Xiangyang Ji, and Simon S Du · 2009
Earlier work this paper cites.
Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert J Kappen · 2013
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Cited alongside, same era.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Unifying PAC and regret: Uniform PAC bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Cited alongside, same era.
Open problem: The dependence of sample complexity lower bounds on planning horizon
Policy certificates: Towards accountable reinforcement learning
Christoph Dann, Lihong Li, Wei Wei, and Emma Brunskill · 2019
Later among the works it cites.
Provably efficient RL with rich observations via latent state decoding
Simon S Du, Akshay Krishnamurthy, Nan Jiang, Alekh Agarwal, Miroslav Dudík, and John Langford · 2019
Later among the works it cites.
Provably efficient maximum entropy exploration
Elad Hazan, Sham Kakade, Karan Singh, and Abby Van Soest · 2019
Later among the works it cites.
Reinforcement learning with convex constraints
Sobhan Miryoosefi, Kianté Brantley, Hal Daume III, Miro Dudik, and Robert E Schapire · 2019
Later among the works it cites.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Andrea Zanette and Emma Brunskill · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nan Jiang and Alekh Agarwal · 2018
Cited alongside, same era.
Is Q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Cited alongside, same era.
Near-optimal time and sample complexities for solving Markov decision processes with a generative model
Aaron Sidford, Mengdi Wang, Xian Wu, Lin Yang, and Yinyu Ye · 2018
Cited alongside, same era.
Reward constrained policy optimization
Chen Tessler, Daniel J Mankowitz, and Shie Mannor · 2018
Cited alongside, same era.
Model-based reinforcement learning with a generative model is minimax optimal
Alekh Agarwal, Sham Kakade, and Lin F Yang · 2019
Cited alongside, same era.
Chi Jin, Akshay Krishnamurthy, Max Simchowitz, and Tiancheng Yu · 2020
Closest in time.
Adaptive reward-free exploration
Emilie Kaufmann, Pierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Edouard Leurent, and Michal Valko · 2020
Closest in time.
Breaking the sample size barrier in model-based reinforcement learning with a generative model
Gen Li, Yuting Wei, Yuejie Chi, Yuantao Gu, and Yuxin Chen · 2020
Closest in time.
Fast active learning for pure exploration in reinforcement learning
Pierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Emilie Kaufmann, Edouard Leurent, and Michal Valko · 2020
Closest in time.
Provably efficient reward-agnostic navigation with linear value iteration
Andrea Zanette, Alessandro Lazaric, Mykel J Kochenderfer, and Emma Brunskill · 2020
Closest in time.