Fetching the paper…
Reading the bibliography…
Efficient exploration in complex environments remains a major challenge for reinforcement learning.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W.R. Thompson · 1933
Earlier work this paper cites.
Some asymptotic theory for the bootstrap
Peter J Bickel and David A Freedman · 1981
Earlier work this paper cites.
The jackknife, the bootstrap and other resampling plans
Bradley Efron · 1982
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
An introduction to the bootstrap
Bradley Efron and Robert J Tibshirani · 1994
Earlier work this paper cites.
Temporal difference learning and td-gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard Sutton and Andrew Barto · 1998
Earlier work this paper cites.
A bayesian framework for reinforcement learning
Malcolm J. A. Strens · 2000
Earlier work this paper cites.
On the Sample Complexity of Reinforcement Learning
Sham Kakade · 2003
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2012
Cited alongside, same era.
Efficient bayes-adaptive reinforcement learning using sample-based search
Arthur Guez, David Silver, and Peter Dayan · 2012
Cited alongside, same era.
Bootstrapping data arrays of arbitrary order
Art B Owen, Dean Eckles, et al · 2012
Cited alongside, same era.
(More) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Cited alongside, same era.
Efficient exploration and value function generalization in deterministic systems
Weight uncertainty in neural networks
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Later among the works it cites.
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Later among the works it cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr et al. Mnih · 2015
Later among the works it cites.
Bootstrapped thompson sampling and deep exploration
Ian Osband and Benjamin Van Roy · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zheng Wen and Benjamin Van Roy · 2013
Cited alongside, same era.
Model-based reinforcement learning and the eluder dimension
Ian Osband and Benjamin Van Roy · 2014
Cited alongside, same era.
Generalization and exploration via randomized value functions
Ian Osband, Benjamin Van Roy, and Zheng Wen · 2014
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2015
Later among the works it cites.
Incentivizing exploration in reinforcement learning with deep predictive models
Bradly C Stadie, Sergey Levine, and Pieter Abbeel · 2015
Later among the works it cites.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2015
Later among the works it cites.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Nando de Freitas, and Marc Lanctot · 2015
Later among the works it cites.