Fetching the paper…
Reading the bibliography…
Posterior sampling for reinforcement learning (PSRL) is an effective method for balancing exploration and exploitation in reinforcement learning.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R · 1933
Earlier work this paper cites.
Improving generalization for temporal difference learning: the successor representation
Dayan, P · 1993
Earlier work this paper cites.
Reinforcement learning: a survey
Kaelbling, L. P., Littman, M. L., and Moore, A. W · 1996
Earlier work this paper cites.
Bayesian Q-Learning
Dearden, R., Friedman, N., and Russell, S. J · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S., Barto, A. G., et al · 1998
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D · 2000
Earlier work this paper cites.
A Bayesian framework for reinforcement learning
Strens, M · 2000
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, P · 2002
Earlier work this paper cites.
Reinforcement Learning and Dynamic Programming Using Function Approximators
Busoniu, L., Babuska, R., Schutter, B. D., and Ernst, D · 2010
Earlier work this paper cites.
(More) efficient reinforcement learning via posterior sampling
Osband, I., Russo, D., and Van Roy, B · 2013
Earlier work this paper cites.
Adam: a method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2015
Cited alongside, same era.
Uncertainty in deep learning
Gal, Y · 2016
Cited alongside, same era.
Deep successor reinforcement learning
Kulkarni, T. D., Saeedi, A., Gautam, S., and Gershman, S. J · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M · 2016
Cited alongside, same era.
Efficient exploration with double uncertain value networks
Moerland, T. M., Broekens, J., and Jonker, C. M · 2017
Later among the works it cites.
Efficient exploration through bayesian deep Q-networks
Azizzadenesheli, K., Brunskill, E., and Anandkumar, A · 2018
Closest in time.
Rainbow: combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M. G., and Silver, D · 2018
Closest in time.
BBQ-Networks: efficient exploration in deep reinforcement learning for task-oriented dialogue systems
Lipton, Z. C., Li, X., Gao, J., Li, L., Ahmed, F., and Deng, L · 2018
Closest in time.
Count-based exploration with the successor representation
Machado, M. C., Bellemare, M. G., and Bowling, M · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Osband, I. and Van Roy, B · 2016
Cited alongside, same era.
Deep reinforcement learning with double Q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., Van Hasselt, H., Lanctot, M., and De Freitas, N · 2016
Cited alongside, same era.
Successor features for transfer in reinforcement learning
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., van Hasselt, H. P., and Silver, D · 2017
Cited alongside, same era.
Shallow updates for deep reinforcement learning
Levine, N., Zahavy, T., Mankowitz, D. J., Tamar, A., and Mannor, S · 2017
Cited alongside, same era.
Eigenoption discovery through the deep successor representation
Machado, M. C., Rosenbaum, C., Guo, X., Liu, M., Tesauro, G., and Campbell, M · 2017
Cited alongside, same era.
Deep exploration via bootstrapped DQN
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B
Cited in the paper.
Closest in time.
The uncertainty Bellman equation and exploration
O’Donoghue, B., Osband, I., Munos, R., and Mnih, V · 2018
Closest in time.
Randomized prior functions for deep reinforcement learning
Osband, I., Aslanides, J., and Cassirer, A · 2018
Closest in time.
Parameter space noise for exploration
Plappert, M., Houthooft, R., Dhariwal, P., Sidor, S., Chen, R. Y., Chen, X., Asfour, T., Abbeel, P., and Andrychowicz, M · 2018
Closest in time.
Randomized value functions via multiplicative normalizing flows
Touati, A., Satija, H., Romoff, J., Pineau, J., and Vincent, P · 2018
Closest in time.
Go-explore: a new approach for hard-exploration problems
Ecoffet, A., Huizinga, J., Lehman, J., Stanley, K. O., and Clune, J · 2019
Closest in time.