Fetching the paper…
Reading the bibliography…
Exploration is a fundamental aspect of Reinforcement Learning, typically implemented using stochastic action-selection.
Efficient memory-based learning for robot control
Andrew William Moore · 1990
Earlier work this paper cites.
Curious model-building control systems
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Efficient exploration in reinforcement learning, 1992
Sebastian B. Thrun · 1992
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
On-line Q-learning using connectionist systems
Gavin A Rummery and Mahesan Niranjan · 1994
Earlier work this paper cites.
Reinforcement driven information acquisition in non-deterministic environments
Jan Storck, Sepp Hochreiter, and Jürgen Schmidhuber · 1995
Earlier work this paper cites.
Reinforcement learning: A survey
Leslie Pack Kaelbling, Michael L Littman, and Andrew W Moore · 1996
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Exploration of multi-state environments: Local measures and back-propagation of uncertainty
Nicolas Meuleau and Paul Bourgine · 1999
Cited alongside, same era.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Cited alongside, same era.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade et al · 2003
Cited alongside, same era.
Pac model-free reinforcement learning
Alexander L Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L Littman · 2006
Cited alongside, same era.
A comprehensive survey of multiagent reinforcement learning
Lucian Busoniu, Robert Babuska, and Bart De Schutter · 2008
Cited alongside, same era.
An analysis of model-based interval estimation for markov decision processes
Alexander L Strehl and Michael L Littman · 2008
Cited alongside, same era.
Exploring compact reinforcement-learning representations with linear regression
Thomas J Walsh, István Szita, Carlos Diuk, and Michael L Littman · 2009
Later among the works it cites.
Planning to be surprised: Optimal bayesian exploration in dynamic environments
Yi Sun, Faustino Gomez, and Jürgen Schmidhuber · 2011
Later among the works it cites.
Value-difference based exploration: adaptive control between epsilon-greedy and softmax
Michel Tokic and Günther Palm · 2011
Later among the works it cites.
Efficient bayes-adaptive reinforcement learning using sample-based search
Arthur Guez, David Silver, and Peter Dayan · 2012
Later among the works it cites.
Learning and exploration in action-perception loops
Daniel Y Little and Friedrich T Sommer · 2014
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Near-bayesian exploration in polynomial time
J Zico Kolter and Andrew Y Ng · 2009
Cited alongside, same era.
Multi-resolution exploration in continuous spaces
Ali Nouri and Michael L Littman · 2009
Cited alongside, same era.
Later among the works it cites.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Later among the works it cites.
Vime: Variational information maximizing exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Later among the works it cites.