Fetching the paper…
Reading the bibliography…
We study an exploration method for model-free RL that generalizes the counter-based exploration bonus methods and takes into account long term exploratory value of actions rather than a single step look-ahead.
Convergence results for single-step on-policy reinforcement-learning algorithms
Singh, Satinder, Jaakkola, Tommi, Littman, Michael L, and Szepesvári, Csaba · 2000
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, Ronen I and Tennenholtz, Moshe · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, Michael and Singh, Satinder · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, Sham Machandranath · 2003
Earlier work this paper cites.
A theoretical analysis of model-based interval estimation
Strehl, Alexander L and Littman, Michael L · 2005
Earlier work this paper cites.
Pac model-free reinforcement learning
Strehl, Alexander L, Li, Lihong, Wiewiora, Eric, Langford, John, and Littman, Michael L · 2006
Earlier work this paper cites.
Probably approximately correct (PAC) exploration in reinforcement learning
Strehl, Alexander L · 2007
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Strehl, Alexander L and Littman, Michael L · 2008
Cited alongside, same era.
The many faces of optimism: a unifying approach
Szita, István and Lőrincz, András · 2008
Cited alongside, same era.
Near-bayesian exploration in polynomial time
Kolter, J Zico and Ng, Andrew Y · 2009
Cited alongside, same era.
A unifying framework for computational reinforcement learning theory
Li, Lihong · 2009
Cited alongside, same era.
Reinforcement learning in finite mdps: Pac analysis
Strehl, Alexander L, Li, Lihong, and Littman, Michael L · 2009
Cited alongside, same era.
Efficient exploration in reinforcement learning
Thrun, Sebastian B · 2009
Learning and exploration in action-perception loops
Little, Daniel Ying-Jeh and Sommer, Friedrich Tobias · 2013
Later among the works it cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, Marc, Srinivasan, Sriram, Ostrovski, Georg, Schaul, Tom, Saxton, David, and Munos, Remi · 2016
Later among the works it cites.
Generalization and exploration via randomized value functions
Osband, Ian, Van Roy, Benjamin, and Wen, Zheng · 2016
Later among the works it cites.
Count-based exploration with neural density models
Ostrovski, Georg, Bellemare, Marc G, Oord, Aaron van den, and Munos, Rémi · 2017
Later among the works it cites.
# exploration: A study of count-based exploration for deep reinforcement learning
Tang, Haoran, Houthooft, Rein, Foote, Davis, Stooke, Adam, Chen, OpenAI Xi, Duan, Yan, Schulman, John, DeTurck, Filip, and Abbeel, Pieter · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Model-based reinforcement learning with nearly tight exploration complexity bounds
Szita, István and Szepesvári, Csaba · 2010
Cited alongside, same era.
Dora the explorer: Directed outreaching reinforcement action-selection
Choshen, Leshem, Fox, Lior, and Loewenstein, Yonatan · 2018
Closest in time.