Fetching the paper…
Reading the bibliography…
We introduce exploration potential, a quantity that measures how much a reinforcement learning agent has explored its environment class.
Model based Bayesian exploration
Richard Dearden, Nir Friedman, and David Andre · 1999
Earlier work this paper cites.
A Bayesian framework for reinforcement learning
Malcolm Strens · 2000
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Universal Artificial Intelligence
Marcus Hutter · 2005
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Jürgen Schmidhuber · 2010
Earlier work this paper cites.
Planning to be surprised: Optimal bayesian exploration in dynamic environments
Yi Sun, Faustino Gomez, and Jürgen Schmidhuber · 2011
Cited alongside, same era.
Universal knowledge-seeking agents for stochastic environments
Laurent Orseau, Tor Lattimore, and Marcus Hutter · 2013
Cited alongside, same era.
Learning to optimize via information-directed sampling
Dan Russo and Benjamin Van Roy · 2014
Cited alongside, same era.
Optimally confident UCB: Improved regret for finite-armed bandits
Tor Lattimore · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Curiosity-driven exploration in deep reinforcement learning via bayesian neural networks
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Closest in time.
Nonparametric General Reinforcement Learning
Jan Leike · 2016
Closest in time.
Learning purposeful behaviour in the absence of rewards
Marlos C Machado and Michael Bowling · 2016
Closest in time.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Closest in time.
Infomax strategies for an optimal balance between exploration and exploitation
Gautam Reddy, Antonio Celani, and Massimo Vergassola · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Marc G Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Cited alongside, same era.