Fetching the paper…
Reading the bibliography…
Although exploration in reinforcement learning is well understood from a theoretical point of view, provably correct methods remain impractical.
Aggregation in dynamic programming
James C Bean, John R Birge, and Robert L Smith · 1987
Earlier work this paper cites.
Eigenvalue bounds on convergence to stationarity for nonreversible markov chains
J. A. Fill · 1991
Earlier work this paper cites.
Efficient reinforcement learning in factored MDPs
Michael Kearns and Daphne Koller · 1999
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Bounded-parameter Markov decision processes
Robert Givan, Sonia Leach, and Thomas Dean · 2000
Earlier work this paper cites.
State abstraction for programmable reinforcement learning agents
David Andre and Stuart J Russell · 2002
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Approximate equivalence of Markov decision processes
Eyal Even-Dar and Yishay Mansour · 2003
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade et al · 2003
Earlier work this paper cites.
Metrics for finite Markov decision processes
Norman Ferns, Prakash Panangaden, and Doina Precup · 2004
Earlier work this paper cites.
Approximate homomorphisms: A framework for non-exact minimization in markov decision processes
Balaraman Ravindran and Andrew G Barto · 2004
Earlier work this paper cites.
Methods for computing state similarity in Markov decision processes
Norman Ferns, Pablo Samuel Castro, Doina Precup, and Prakash Panangaden · 2006
Earlier work this paper cites.
Towards a unified theory of state abstraction for MDPs
Lihong Li, Thomas J Walsh, and Michael L Littman · 2006
Cited alongside, same era.
PAC model-free reinforcement learning
Alexander L Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L Littman · 2006
Cited alongside, same era.
Pseudometrics for state aggregation in average reward Markov decision processes
Ronald Ortner · 2007
Cited alongside, same era.
An analysis of model-based interval estimation for Markov decision processes
Alexander L Strehl and Michael L Littman · 2008
Cited alongside, same era.
Near-Bayesian exploration in polynomial time
J Zico Kolter and Andrew Y Ng · 2009
Cited alongside, same era.
A unifying framework for computational reinforcement learning theory
Lihong Li · 2009
Cited alongside, same era.
Near optimal behavior via approximate state abstraction
David Abel, D. Ellis Hershkowitz, and Michael L. Littman · 2016
Later among the works it cites.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Later among the works it cites.
Jan Leike · 2016
Later among the works it cites.
Minimax regret bounds for Reinforcement Learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Later among the works it cites.
Unifying PAC and Regret: Uniform PAC bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Later among the works it cites.
Exploration–Exploitation in MDPs with Options
Ronan Fruit and Alessandro Lazaric · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Cited alongside, same era.
Model-based reinforcement learning with nearly tight exploration complexity bounds
István Szita and Csaba Szepesvári · 2010
Cited alongside, same era.
On the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert Kappen · 2012
Cited alongside, same era.
Adaptive aggregation for reinforcement learning in average reward markov decision processes
Ronald Ortner · 2013
Cited alongside, same era.
PAC-inspired option discovery in lifelong reinforcementlearning
Emma Brunskill and Lihong Li · 2014
Cited alongside, same era.
Extreme state aggregation beyond MDPs
Marcus Hutter · 2014
Cited alongside, same era.
Later among the works it cites.
Count-based exploration with Neural Density Models
Georg Ostrovski, Marc G. Bellemare, Aäron van den Oord, and Rémi Munos · 2017
Later among the works it cites.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell · 2017
Later among the works it cites.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2018
Closest in time.
Noisy networks for exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Ian Osband, Alex Graves, Vlad Mnih, Remi Munos, Demis Hassabis, Olivier Pietquin, et al · 2018
Closest in time.
Efficient bias-span-constrained exploration-exploitation in reinforcement learning
Ronan Fruit, Matteo Pirotta, Alessandro Lazaric, and Ronald Ortner · 2018
Closest in time.
Exploration in structured reinforcement learning
Jungseul Ok, Alexandre Proutiere, and Damianos Tranos · 2018
Closest in time.
Parameter space noise for exploration
Matthias Plappert, Rein Houthooft, Prafulla Dhariwal, Szymon Sidor, Richard Y. Chen, Xi Chen, Tamim Asfour, Pieter Abbeel, and Marcin Andrychowicz · 2018
Closest in time.