Fetching the paper…
Reading the bibliography…
While off-policy temporal difference (TD) methods have widely been used in reinforcement learning due to their efficiency and simple implementation, their Bayesian counterparts have not been utilized as frequently.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. R. Thompson · 1933
Earlier work this paper cites.
Stochastic models, estimation and control
P. S. Maybeck · 1982
Earlier work this paper cites.
Q-learning
C. J. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Reinforcement learning: A survey
L. P. Kaelbling, M. L. Littman, and A. W. Moore · 1996
Earlier work this paper cites.
Tractable inference for complex stochastic processes
X. Boyen and D. Koller · 1998
Earlier work this paper cites.
Bayesian q-learning
R. Dearden, N. Friedman, and S. Russell · 1998
Earlier work this paper cites.
Model based bayesian exploration
R. Dearden, N. Friedman, and D. Andre · 1999
Earlier work this paper cites.
A bayesian approach to online learning
M. Opper · 1999
Earlier work this paper cites.
A bayesian framework for reinforcement learning
M. Strens · 2000
Earlier work this paper cites.
Optimal learning: Computational procedures for bayes-adaptive markov decision processes
M. Duff · 2002
Earlier work this paper cites.
On the convergence of optimistic policy iteration
J. N. Tsitsiklis · 2002
Earlier work this paper cites.
Bayes meets bellman: The gaussian process approach to temporal difference learning
Y. Engel, S. Mannor, and R. Meir · 2003
Earlier work this paper cites.
Least-squares policy iteration
M. G. Lagoudakis and R. Parr · 2003
Cited alongside, same era.
Reinforcement learning with gaussian processes
Y. Engel, S. Mannor, and R. Meir · 2005
Cited alongside, same era.
Bayesian policy gradient algorithm
M. Ghavamzadeh and Y. Engel · 2006
Cited alongside, same era.
An analytic solution to discrete bayesian reinforcement learning
P. Poupart, N. Vlassis, J. Hoey, and K. Regan · 2006
Cited alongside, same era.
Kalman temporal differences
M. Geist and O. Pietquin · 2010
Cited alongside, same era.
Double q-learning
H. V. Hasselt · 2010
Cited alongside, same era.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Generalization and exploration via randomized value functions
I. Osband, B. V. Roy, and Z. Wen · 2014
Later among the works it cites.
Bayesian reinforcement learning: A survey
M. Ghavamzadeh, S. Mannor, J. Pineau, and A. Tamar · 2015
Later among the works it cites.
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2015
Later among the works it cites.
Openai gym, 2016
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Later among the works it cites.
Q ( λ \lambda ) with off-policy corrections
A. Harutyunyan, M. G. Bellemare, T. Stepleton, and R. Munos · 2016
Later among the works it cites.
Deep reinforcement learning with double q-learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. D. Ziebart · 2010
Cited alongside, same era.
Efficient bayes-adaptive reinforcement learning using sample-based search
A. Guez, D. Silver, and P. Dayan · 2012
Cited alongside, same era.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Cited alongside, same era.
(more) efficient reinforcement learning via posterior sampling
I. Osband, D. Russo, and B. V. Roy · 2013
Cited alongside, same era.
Off-policy reinforcement learning with gaussian process
G. Chowdhary, M. Liu, R. Grande, T. Walsh, J. How, and L. Carin · 2014
Cited alongside, same era.
H. V. Hasselt, A. Guez, and D. Silver · 2016
Later among the works it cites.
Deep exploration via bootstrapped dqn
I. Osband, C. Blundell, A. Pritzel, and B. V. Roy · 2016
Later among the works it cites.
A distributional perspective on reinforcement learning
M. G. Bellemare, W. Dabney, , and R. Munos · 2017
Closest in time.
Openai baselines
P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, and Y. Wu · 2017
Closest in time.
The uncertainty bellman equation and exploration
B. O’Donoghue, I. Osband, R. Munos, and V. Mnih · 2017
Closest in time.
Efficient exploration through bayesian deep q-networks
K. Azizzadenesheli, E. Brunskill, and A. Anandkumar · 2018
Closest in time.