Fetching the paper…
Reading the bibliography…
This paper studies a recent proposal to use randomized value functions to drive exploration in reinforcement learning.
A bayesian framework for reinforcement learning
Malcolm Strens · 2000
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade et al · 2003
Earlier work this paper cites.
Least-squares policy iteration
Michail G Lagoudakis and Ronald Parr · 2003
Earlier work this paper cites.
Inequalities for the l1 deviation of the empirical distribution
Tsachy Weissman, Erik Ordentlich, Gadiel Seroussi, Sergio Verdu, and Marcelo J Weinberger · 2003
Earlier work this paper cites.
Pac model-free reinforcement learning
Alexander L Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L Littman · 2006
Earlier work this paper cites.
A bayesian sampling approach to exploration in reinforcement learning
John Asmuth, Lihong Li, Michael L Littman, Ali Nouri, and David Wingate · 2009
Earlier work this paper cites.
Reinforcement learning in finite mdps: Pac analysis
Alexander L Strehl, Lihong Li, and Michael L Littman · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
Shipra Agrawal and Navin Goyal · 2013
Cited alongside, same era.
Concentration inequalities: A nonasymptotic theory of independence
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart · 2013
Cited alongside, same era.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Cited alongside, same era.
Linear thompson sampling revisited
Marc Abeille, Alessandro Lazaric, et al · 2017
Cited alongside, same era.
Deep exploration via randomized value functions
Ian Osband, Benjamin Van Roy, Daniel Russo, and Zheng Wen · 2017
Later among the works it cites.
Efficient exploration through bayesian deep q-networks
Kamyar Azizzadenesheli, Emma Brunskill, and Animashree Anandkumar · 2018
Later among the works it cites.
Noisy networks for exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Ian Osband, Alex Graves, Vlad Mnih, Remi Munos, Demis Hassabis, Olivier Pietquin, et al · 2018
Later among the works it cites.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Later among the works it cites.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2018
Later among the works it cites.
Randomized prior functions for deep reinforcement learning
Ian Osband, John Aslanides, and Albin Cassirer · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Shipra Agrawal and Randy Jia · 2017
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Cited alongside, same era.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy
Cited in the paper.
Generalization and exploration via randomized value functions
Ian Osband, Benjamin Van Roy, and Zheng Wen
Cited in the paper.
Later among the works it cites.
Randomized value functions via multiplicative normalizing flows
Ahmed Touati, Harsh Satija, Joshua Romoff, Joelle Pineau, and Pascal Vincent · 2018
Later among the works it cites.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2019
Closest in time.
Randomised bayesian least-squares policy iteration
Nikolaos Tziortziotis, Christos Dimitrakakis, and Michalis Vazirgiannis · 2019
Closest in time.