Fetching the paper…
Reading the bibliography…
We propose a framework based on distributional reinforcement learning and recent attempts to combine Bayesian parameter updates with deep reinforcement learning.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R. Thompson · 1933
Earlier work this paper cites.
Feature based methods for large scale dynamic programming
John N. Tsitsiklis and Benjamin Van Roy · 1996
Earlier work this paper cites.
Bayesian q learning
Richard Dearden, Nir Friedman, and Stuart Russel · 1998
Earlier work this paper cites.
Nonparametric return distribution approximation for reinforcement learning
Tetsuro Morimura, Masashi Sugiyama, Hisashi Kashima, Hirotaka Hachiya, and Toshiyuki Tanaka · 2010
Earlier work this paper cites.
Parametric return density estimation for reinforcement learning
Tetsuro Morimura, Masashi Sugiyama, Hisashi Kashima, Hirotaka Hachiya, and Toshiyuki Tanaka · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Earlier work this paper cites.
Black box variational inference
Rejesh Ranganath, Sean Gerrish, and David M. Blei · 2014
Earlier work this paper cites.
Bootstrapped thompson sampling and deep exploration
Ian Osband and Benjamin Van Roy · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Philipp Moritz, Michael I. Jordan, and Pieter Abbeel · 2015
Cited alongside, same era.
Openai gym
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Cited alongside, same era.
End to end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Cited alongside, same era.
Efficient exploration through bayesian deep q networks
Kamyar Azizzadenesheli, Emma Brunskill, and Animashree Anandkumar · 2017
Later among the works it cites.
A distributional perspective on reinforcement learning
Marc G. Bellemare, Will Dabney, and Remi Munos · 2017
Later among the works it cites.
Variational inference: A review for statisticians
David M. Blei, Alp Kucukelbir, and Jon D. McAuliffe · 2017
Later among the works it cites.
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Ian Osband, Alex Graves, Vlad Mnih, Remi Munos, Demis Hassabis, Ilivier Pietquin, Charles Blundell, and Shane Legg · 2017
Later among the works it cites.
Bayesian policy gradients via alpha divergence dropout inference
Peter Henderson, Thang Doan, Riashat Islam, and David Meger · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Efficient dialogue policy learning with bbq-networks
Zachary C. Lipton, Xiujun Li, Jianfeng Gao, Lihong Li, Faisal Ahmed, and Li Deng · 2016
Cited alongside, same era.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Cited alongside, same era.
Parameter space noise for exploration
Matthias Plappert, Rein Houthooft, Prafulla Dhariwal, Szymon Sidor, Richard Y. Chen, Xi Chen, Tamim Asfour, Pieter Abbeel, and Marcin Andrychowicz · 2016
Cited alongside, same era.
Thomas M. Moerland, Joost Broekens, and Catholijn M. Jonker · 2017
Later among the works it cites.
Deep exploration via randomized value functions
Ian Osband, daniel Russo, Zheng Wen, and Benjamin Van Roy · 2017
Later among the works it cites.
Tutorial on thompson sampling
Daniel Russo · 2017
Later among the works it cites.
Variational deep q network
Yunhao Tang and Alp Kucukelbir · 2017
Later among the works it cites.