Fetching the paper…
Reading the bibliography…
We study the use of randomized value functions to guide deep exploration in reinforcement learning.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
Discounted dynamic programming
David Blackwell · 1965
Earlier work this paper cites.
Rules for ordering uncertain prospects
Josef Hadar and William R Russell · 1969
Earlier work this paper cites.
The efficiency analysis of choices involving risk
G Hanoch and H Levy · 1969
Earlier work this paper cites.
Some asymptotic theory for the bootstrap
Peter J Bickel and David A Freedman · 1981
Earlier work this paper cites.
The jackknife, the bootstrap and other resampling plans , volume 38
Bradley Efron · 1982
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
An introduction to the bootstrap
Bradley Efron and Robert J Tibshirani · 1994
Earlier work this paper cites.
Temporal difference learning and TD-gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Neuro-Dynamic Programming
Dimitri P. Bertsekas and John Tsitsiklis · 1996
Earlier work this paper cites.
Increasing risk: some direct constructions
Mark Machina and John Pratt · 1997
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
John N Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
Bayesian Q-learning
Richard Dearden, Nir Friedman, and Stuart J. Russell · 1998
Earlier work this paper cites.
Efficient reinforcement learning in factored MDPs
Michael J. Kearns and Daphne Koller · 1999
Earlier work this paper cites.
Off-policy temporal-difference learning with function approximation
Doina Precup, Richard Sutton, and Sanjoy Dasgupta · 2001
Earlier work this paper cites.
R-max - a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I. Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael J. Kearns and Satinder P. Singh · 2002
Earlier work this paper cites.
On the Sample Complexity of Reinforcement Learning
Sham Kakade · 2003
Earlier work this paper cites.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Bootstrap prediction and bayesian prediction under misspecified models
Tadayoshi Fushiki · 2005
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
Peter Auer and Ronald Ortner · 2006
Earlier work this paper cites.
PAC model-free reinforcement learning
Alexander L. Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L. Littman · 2006
Earlier work this paper cites.
Probably approximately correct (PAC) exploration in reinforcement learning
Alexander L Strehl · 2007
Earlier work this paper cites.
Knows what it knows: a framework for self-aware learning
Lihong Li, Michael L Littman, and Thomas J Walsh · 2008
Earlier work this paper cites.
Natural evolution strategies
Daan Wierstra, Tom Schaul, Jan Peters, and Juergen Schmidhuber · 2008
Earlier work this paper cites.
REGAL: A regularization based algorithm for reinforcement learning in weakly communicating MDPs
Peter L. Bartlett and Ambuj Tewari · 2009
Cited alongside, same era.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Richard Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári, and Eric Wiewiora · 2009
Cited alongside, same era.
Transfer learning for reinforcement learning domains: A survey
Matthew E Taylor and Peter Stone · 2009
Cited alongside, same era.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Cited alongside, same era.
Reducing reinforcement learning to kwik online regression
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2015
Later among the works it cites.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Later among the works it cites.
Angrier birds: Bayesian reinforcement learning
Imanol Arrieta Ibarra, Bernardo Ramos, and Lars Roemheld · 2016
Later among the works it cites.
Safe and efficient off-policy reinforcement learning
Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc Bellemare · 2016
Later among the works it cites.
Deep Exploration via Randomized Value Functions
Ian Osband · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lihong Li and Michael L Littman · 2010
Cited alongside, same era.
Algorithms for Reinforcement Learning
Csaba Szepesvári · 2010
Cited alongside, same era.
Regret bounds for the adaptive control of linear quadratic systems
Yasin Abbasi-Yadkori and Csaba Szepesvári · 2011
Cited alongside, same era.
Optimal Learning
Warren Powell and Ilya Ryzhov · 2011
Cited alongside, same era.
Experience replay for real-time reinforcement learning control
Sander Adam, Lucian Busoniu, and Robert Babuska · 2012
Cited alongside, same era.
Analysis of Thompson sampling for the multi-armed bandit problem
Shipra Agrawal and Navin Goyal · 2012
Cited alongside, same era.
Efficient reinforcement learning for high dimensional linear quadratic systems
Morteza Ibrahimi, Adel Javanmard, and Benjamin V Roy · 2012
Cited alongside, same era.
Ian Osband and Benjamin Van Roy · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Later among the works it cites.
#Exploration: A study of count-based exploration for deep reinforcement learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Later among the works it cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Closest in time.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Closest in time.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Closest in time.
Ensemble sampling
Xiuyuan Lu and Benjamin Van Roy · 2017
Closest in time.
The uncertainty Bellman equation and exploration
Brendan O’Donoghue, Ian Osband, Remi Munos, and Volodymyr Mnih · 2017
Closest in time.
Why is posterior sampling better than optimism for reinforcement learning?
Ian Osband and Benjamin Van Roy · 2017
Closest in time.
Parameter Space Noise for Exploration in Deep Reinforcement Learning
Matthias Plappert · 2017
Closest in time.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Closest in time.
Efficient exploration through Bayesian deep q-networks
Kamyar Azizzadenesheli, Emma Brunskill, and Animashree Anandkumar · 2018
Closest in time.
Noisy networks for exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Matteo Hessel, Ian Osband, Alex Graves, Volodymyr Mnih, Remi Munos, Demis Hassabis, Olivier Pietquin, Charles Blundell, and Shane Legg · 2018
Closest in time.
BBQ-networks: Efficient exploration in deep reinforcement learning for task-oriented dialogue systems
Zachary Lipton, Xiujun Li, Jianfeng Gao, Lihong Li, Faisal Ahmed, and Li Deng · 2018
Closest in time.
Randomized prior functions for deep reinforcement learning
Ian Osband, John Aslanides, and Albin Cassirer · 2018
Closest in time.
A tutorial on Thompson sampling
Daniel J Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, Zheng Wen, et al · 2018
Closest in time.
Reinforcement Learning: An Introduction, Second Edition
Richard Sutton and Andrew Barto · 2018
Closest in time.
Bootstrap thompson sampling and sequential decision problems in the behavioral sciences
Dean Eckles and Maurits Kaptein · 2019
Closest in time.
Worst-case regret bounds for exploration via randomized value functions
Daniel Russo · 2019
Closest in time.