Fetching the paper…
Reading the bibliography…
This paper studies directed exploration for reinforcement learning agents by tracking uncertainty about the value of each available action.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R. (1933) · 1933
Earlier work this paper cites.
The variance of discounted Markov decision processes
Sobel, M. J. (1982) · 1982
Earlier work this paper cites.
Mean, variance, and probabilistic criteria in finite Markov decision processes: a review
White, D. (1988) · 1988
Earlier work this paper cites.
Bayesian Q-learning
Dearden, R., Friedman, N., and Russell, S. (1998) · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Model based Bayesian exploration
Dearden, R., Friedman, N., and Andre, D. (1999) · 1999
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., and Fischer, P. (2002) · 2002
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, R. I. and Tennenholtz, M. (2002) · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S. (2002) · 2002
Earlier work this paper cites.
Bayes meets Bellman: The Gaussian process approach to temporal difference learning
Engel, Y., Mannor, S., and Meir, R. (2003) · 2003
Earlier work this paper cites.
Information theory, inference and learning algorithms
MacKay, D. J. (2003) · 2003
Earlier work this paper cites.
Gaussian Processes in Reinforcement Learning
Rasmussen, C. E., Kuss, M., et al. (2003) · 2003
Earlier work this paper cites.
Bias and variance in value function estimation
Mannor, S., Simester, D., Sun, P., and Tsitsiklis, J. N. (2004) · 2004
Earlier work this paper cites.
Reinforcement learning with Gaussian processes
Engel, Y., Mannor, S., and Meir, R. (2005) · 2005
Cited alongside, same era.
Bandit based monte-carlo planning
Kocsis, L. and Szepesvári, C. (2006) · 2006
Cited alongside, same era.
Bias and variance approximation in value function estimates
Mannor, S., Simester, D., Sun, P., and Tsitsiklis, J. N. (2007) · 2007
Cited alongside, same era.
An empirical evaluation of thompson sampling
Chapelle, O. and Li, L. (2011) · 2011
Cited alongside, same era.
PILCO: A model-based and data-efficient approach to policy search
Deisenroth, M. and Rasmussen, C. E. (2011) · 2011
Cited alongside, same era.
Mean-variance optimization in Markov decision processes
Mannor, S. and Tsitsiklis, J. (2011) · 2011
Cited alongside, same era.
Uncertainty in deep learning
Gal, Y. (2016) · 2016
Later among the works it cites.
Dropout as a Bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z. (2016) · 2016
Later among the works it cites.
Improving PILCO with bayesian neural network dynamics models
Gal, Y., McAllister, R. T., and Rasmussen, C. E. (2016) · 2016
Later among the works it cites.
Vime: Variational information maximizing exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., De Turck, F., and Abbeel, P. (2016) · 2016
Later among the works it cites.
Deep exploration via bootstrapped DQN
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B. (2016) · 2016
Later among the works it cites.
Learning the variance of the reward-to-go
Tamar, A., Di Castro, D., and Mannor, S. (2016) · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bayesian learning via stochastic gradient Langevin dynamics
Welling, M. and Teh, Y. W. (2011) · 2011
Cited alongside, same era.
Efficient Bayes-adaptive reinforcement learning using sample-based search
Guez, A., Silver, D., and Dayan, P. (2012) · 2012
Cited alongside, same era.
Parametric return density estimation for reinforcement learning
Morimura, T., Sugiyama, M., Kashima, H., Hachiya, H., and Tanaka, T. (2012) · 2012
Cited alongside, same era.
Generalization and exploration via randomized value functions
Osband, I., Van Roy, B., and Wen, Z. (2014) · 2014
Cited alongside, same era.
Weight uncertainty in neural networks
Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D. (2015) · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, B. C., Levine, S., and Abbeel, P. (2015) · 2015
Cited alongside, same era.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R. (2017) · 2017
Closest in time.
Uncertainty Decomposition in Bayesian Neural Networks with Latent Variables
Depeweg, S., Hernández-Lobato, J. M., Doshi-Velez, F., and Udluft, S. (2017) · 2017
Closest in time.
Parameter Space Noise for Exploration
Matthias, P., Rein, H., Prafulla, D., Szymon, S., Richard Y., C., Xi, C., Tamim, A., Pieter, A., and Marcin, A. (2017) · 2017
Closest in time.
Learning Multimodal Transition Dynamics for Model-Based Reinforcement Learning
Moerland, T. M., Broekens, J., and Jonker, C. M. (2017) · 2017
Closest in time.
The Uncertainty Bellman Equation and Exploration
O’Donoghue, B., Osband, I., Munos, R., and Mnih, V. (2017) · 2017
Closest in time.
Generalized exploration in policy search
van Hoof, H., Tanneberg, D., and Peters, J. (2017) · 2017
Closest in time.