Fetching the paper…
Reading the bibliography…
This paper studies the potential of the return distribution for exploration in deterministic reinforcement learning (RL) environments.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, William R · 1933
Earlier work this paper cites.
The variance of discounted Markov decision processes
Sobel, Matthew J · 1982
Earlier work this paper cites.
Mean, variance, and probabilistic criteria in finite Markov decision processes: a review
White, DJ · 1988
Earlier work this paper cites.
Bayesian Q-learning
Dearden, Richard, Friedman, Nir, and Russell, Stuart · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Sutton, Richard S and Barto, Andrew G · 1998
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, Peter, Cesa-Bianchi, Nicolo, and Fischer, Paul · 2002
Earlier work this paper cites.
Reinforcement learning with Gaussian processes
Engel, Yaakov, Mannor, Shie, and Meir, Ron · 2005
Earlier work this paper cites.
The matrix cookbook
Petersen, Kaare Brandt, Pedersen, Michael Syskind, et al · 2008
Earlier work this paper cites.
Mean-variance optimization in Markov decision processes
Mannor, Shie and Tsitsiklis, John · 2011
Earlier work this paper cites.
Efficient Bayes-adaptive reinforcement learning using sample-based search
Guez, Arthur, Silver, David, and Dayan, Peter · 2012
Earlier work this paper cites.
Parametric return density estimation for reinforcement learning
Morimura, Tetsuro, Sugiyama, Masashi, Kashima, Hisashi, Hachiya, Hirotaka, and Tanaka, Toshiyuki · 2012
Cited alongside, same era.
Bayesian inference in monte-carlo tree search
Tesauro, Gerald, Rajan, VT, and Segal, Richard · 2012
Cited alongside, same era.
Generalization and exploration via randomized value functions
Osband, Ian, Van Roy, Benjamin, and Wen, Zheng · 2014
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Bellemare, Marc, Srinivasan, Sriram, Ostrovski, Georg, Schaul, Tom, Saxton, David, and Munos, Remi · 2016
Cited alongside, same era.
Learning and policy search in stochastic dynamical systems with bayesian neural networks
Bayesian Policy Gradients via Alpha Divergence Dropout Inference
Henderson, Peter, Doan, Thang, Islam, Riashat, and Meger, David · 2017
Later among the works it cites.
Bayesian Q-learning with Assumed Density Filtering
Jeong, Heejin and Lee, Daniel D · 2017
Later among the works it cites.
Monte-carlo tree search by best arm identification
Kaufmann, Emilie and Koolen, Wouter M · 2017
Later among the works it cites.
Efficient exploration with Double Uncertain Value Networks
Moerland, Thomas M, Broekens, Joost, and Jonker, Catholijn M · 2017
Later among the works it cites.
A Tutorial on Thompson Sampling
Russo, Daniel, Van Roy, Benjamin, Kazerouni, Abbas, and Osband, Ian · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Depeweg, Stefan, Hernández-Lobato, José Miguel, Doshi-Velez, Finale, and Udluft, Steffen · 2016
Cited alongside, same era.
Improving PILCO with bayesian neural network dynamics models
Gal, Yarin, McAllister, Rowan Thomas, and Rasmussen, Carl Edward · 2016
Cited alongside, same era.
Deep exploration via bootstrapped DQN
Osband, Ian, Blundell, Charles, Pritzel, Alexander, and Van Roy, Benjamin · 2016
Cited alongside, same era.
Learning the variance of the reward-to-go
Tamar, Aviv, Di Castro, Dotan, and Mannor, Shie · 2016
Cited alongside, same era.
Efficient Exploration through Bayesian Deep Q-Networks
Azizzadenesheli, Kamyar, Brunskill, Emma, and Anandkumar, Animashree · 2017
Cited alongside, same era.
A distributional perspective on reinforcement learning
Bellemare, Marc G, Dabney, Will, and Munos, Rémi · 2017
Cited alongside, same era.
Learning Multimodal Transition Dynamics for Model-Based Reinforcement Learning
Moerland, Thomas M, Broekens, Joost, and Jonker, Catholijn M
Cited in the paper.
Later among the works it cites.
Tang, Yunhao and Kucukelbir, Alp · 2017
Later among the works it cites.
Monte Carlo Tree Search for Asymmetric Trees
Moerland, Thomas M, Broekens, Joost, Plaat, Aske, and Jonker, Catholijn M · 2018
Closest in time.
Randomized Prior Functions for Deep Reinforcement Learning
Osband, Ian, Aslanides, John, and Cassirer, Albin · 2018
Closest in time.
Exploration by Distributional Reinforcement Learning
Tang, Yunhao and Agrawal, Shipra · 2018
Closest in time.