Fetching the paper…
Reading the bibliography…
Reward-free exploration is a reinforcement learning setting studied by Jin et al.
A possibility for implementing curiosity and boredom in model-building neural controllers
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Efficient reinforcement learning
Claude-Nicolas Fiechter · 1994
Earlier work this paper cites.
Finite-sample convergence rates for q-learning and indirect algorithms
Michael J. Kearns and Satinder P. Singh · 1998
Earlier work this paper cites.
R-MAX - A general polynomial time algorithm for near-optimal reinforcement learning
Ronen I. Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael J. Kearns and Satinder P. Singh · 2002
Earlier work this paper cites.
On the Sample Complexity of Reinforcement Learning
Sham Kakade · 2003
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Nuttapong Chentanez, Andrew G. Barto, and Satinder P. Singh · 2005
Earlier work this paper cites.
Action Elimination and Stopping Conditions for the Multi-Armed Bandit and Reinforcement Learning Problems
E. Even-Dar, S. Mannor, and Y. Mansour · 2006
Earlier work this paper cites.
PAC model-free reinforcement learning
Alexander L. Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L. Littman · 2006
Earlier work this paper cites.
On reward-free reinforcement learning with linear function approximation
Ruosong Wang, Simon S Du, Lin F Yang, and Ruslan Salakhutdinov · 2006
Earlier work this paper cites.
Task-agnostic exploration in reinforcement learning
Xuezhou Zhang, Yuzhe Ma, and Adish Singla · 2006
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Alexander L. Strehl and Michael L. Littman · 2008
Cited alongside, same era.
Optimism in Reinforcement Learning and Kullback-Leibler Divergence
S. Filippi, O. Cappé, and A. Garivier · 2010
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Cited alongside, same era.
On the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Bert Kappen · 2012
Cited alongside, same era.
Autonomous exploration for navigating in mdps
Shiau Hong Lim and Peter Auer · 2012
Cited alongside, same era.
An information-theoretic approach to curiosity-driven reinforcement learning
Susanne Still and Doina Precup · 2012
Cited alongside, same era.
Count-based exploration with neural density models
Georg Ostrovski, Marc G Bellemare, Aäron van den Oord, and Rémi Munos · 2017
Later among the works it cites.
# exploration: A study of count-based exploration for deep reinforcement learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip DeTurck, and Pieter Abbeel · 2017
Later among the works it cites.
Provably efficient maximum entropy exploration
Elad Hazan, Sham M. Kakade, Karan Singh, and Abby Van Soest · 2018
Later among the works it cites.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sébastien Bubeck, and Michael I. Jordan · 2018
Later among the works it cites.
Autonomous exploration for navigating in non-stationary cmps, 2019
Pratik Gajane, Ronald Ortner, Peter Auer, and Csaba Szepesvari · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Cited alongside, same era.
Variational information maximisation for intrinsically motivated reinforcement learning
Shakir Mohamed and Danilo Jimenez Rezende · 2015
Cited alongside, same era.
Information theoretically aided reinforcement learning for embodied agents, 2016
Guido Montufar, Keyan Ghazi-Zahedi, and Nihat Ay · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Cited alongside, same era.
Jean Tarbouriech, Evrard Garcelon, Michal Valko, Matteo Pirotta, and Alessandro Lazaric · 2019
Later among the works it cites.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Andrea Zanette and Emma Brunskill · 2019
Later among the works it cites.
Near-optimal regret bounds for stochastic shortest path
Alon Cohen, Haim Kaplan, Yishay Mansour, and Aviv Rosenberg · 2020
Closest in time.
Reward-free exploration for reinforcement learning
Chi Jin, Akshay Krishnamurthy, Max Simchowitz, and Tiancheng Yu · 2020
Closest in time.
Planning in markov decision processes with gap-dependent sample complexity
Anders Jonsson, Emilie Kaufmann, Pierre Ménard, Omar Darwiche-Domingues, Edouard Leurent, and Michal Valko · 2020
Closest in time.