Fetching the paper…
Reading the bibliography…
Algorithms that tackle deep exploration -- an important challenge in reinforcement learning -- have relied on epistemic uncertainty representation through ensembles or other hypermodels, exploration bonuses, or visitation count distributions.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
Playing Atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
(More) efficient reinforcement learning via posterior sampling
I. Osband, D. Russo, and B. Van Roy · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Stochastic processes and applications: diffusion processes, the Fokker-Planck and Langevin equations , volume 60
Grigorios A Pavliotis · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Deep exploration via bootstrapped DQN
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Cited alongside, same era.
Ensemble sampling
Xiuyuan Lu and Benjamin Van Roy · 2017
Cited alongside, same era.
Sharp convergence rates for Langevin dynamics in the nonconvex setting
Xiang Cheng, Niladri S Chatterji, Yasin Abbasi-Yadkori, Peter L Bartlett, and Michael I Jordan · 2018
Cited alongside, same era.
Noisy networks for exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Matteo Hessel, Ian Osband, Alex Graves, Volodymyr Mnih, R/emi Munos, Demis Hassabis, Olivier Pietquin, Charles Blundell, and Shane Legg · 2018
Cited alongside, same era.
Randomized prior functions for deep reinforcement learning
Ian Osband, John Aslanides, and Albin Cassirer · 2018
Cited alongside, same era.
A tutorial on Thompson sampling
Daniel J Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, Zheng Wen, et al · 2018
Posterior sampling networks
Vikranth R Dwaracherla, Benjamin Van Roy, and Morteza Ibrahimi · 2019
Later among the works it cites.
Deep exploration via randomized value functions
Ian Osband, Daniel Russo, Zheng Wen, and Benjamin Van Roy · 2019
Later among the works it cites.
Hypermodels for exploration
Vikranth Dwaracherla, Xiuyuan Lu, Morteza Ibrahimi, Zheng Osband, Ian Oand Wen, and Benjamin Van Roy · 2020
Closest in time.
On approximate thompson sampling with langevin algorithms
Eric Mazumdar, Aldo Pacchiano, Yian Ma, Michael Jordan, and Peter Bartlett · 2020
Closest in time.
Behaviour suite for reinforcement learning
Ian Osband, Yotam Doron, Matteo Hassel, John Aslanides, Eren Sezener, Andre Saraiva, Katrina Mckinney, Tor Lattimore, Csaba Szepesvari, Satinder Singh, Benjamin Van Roy, Richard Sutton, David Silver, and Hado Van Hasselt · 2020
Closest in time.
Bayesian learning via stochastic gradient Langevin dynamics
Max Welling and Yee W Teh · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Closest in time.