Fetching the paper…
Reading the bibliography…
Efficient exploration is one of the key challenges for reinforcement learning (RL) algorithms.
Optimal strategy for item presentation in a learning process
William Karush and RE Dear · 1967
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Thomas G Dietterich · 2000
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Convergence of optimistic and incremental q-learning
Eyal Even-Dar and Yishay Mansour · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Learning rates for q-learning
Eyal Even-Dar and Yishay Mansour · 2003
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade et al · 2003
Earlier work this paper cites.
Laplacians and the cheeger inequality for directed graphs
Fan Chung · 2005
Earlier work this paper cites.
The diameter and laplacian eigenvalues of directed graphs
Fan Chung · 2006
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
Peter Auer and Ronald Ortner · 2007
Cited alongside, same era.
Proto-value functions: A laplacian framework for learning representation and control in markov decision processes
Sridhar Mahadevan and Mauro Maggioni · 2007
Cited alongside, same era.
The epoch-greedy algorithm for multi-armed bandits with side information
John Langford and Tong Zhang · 2008
Cited alongside, same era.
A unifying framework for computational reinforcement learning theory
Lihong Li · 2009
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Cited alongside, same era.
Cover times, blanket times, and majorizing measures
Jian Ding, James R Lee, and Yuval Peres · 2011
Sample complexity of multi-task reinforcement learning
Emma Brunskill and Lihong Li · 2013
Later among the works it cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Later among the works it cites.
How hard is my mdp?” the distribution-norm to the rescue”
Odalric-Ambrym Maillard, Timothy A Mann, and Shie Mannor · 2014
Later among the works it cites.
Complexity and cooperation in q-learning
Steven D Whitehead · 2014
Later among the works it cites.
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sample complexity bounds of exploration
Lihong Li · 2012
Cited alongside, same era.
Incremental model-based learners with formal learning-time guarantees
Alexander L Strehl, Lihong Li, and Michael L Littman · 2012
Cited alongside, same era.
Nan Jiang, Satinder Singh, and Ambuj Tewari · 2016
Later among the works it cites.
Exploiting the natural exploration in contextual bandits
Hamsa Bastani, Mohsen Bayati, and Khashayar Khosravi · 2017
Later among the works it cites.
Markov chains and mixing times , volume 107
David A Levin and Yuval Peres · 2017
Later among the works it cites.