Fetching the paper…
Reading the bibliography…
We consider the exploration/exploitation problem in reinforcement learning.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R · 1933
Earlier work this paper cites.
Dynamic programming
Bellman, R · 1957
Earlier work this paper cites.
The variance of discounted Markov decision processes
Sobel, M. J · 1982
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, T. L. and Robbins, H · 1985
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Learning from delayed rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
Reinforcement Learning: an Introduction
Sutton, R. and Barto, A · 1998
Earlier work this paper cites.
Least-squares temporal difference learning
Boyan, J. A · 1999
Earlier work this paper cites.
A Bayesian framework for reinforcement learning
Strens, M · 2000
Earlier work this paper cites.
R-max: A general polynomial time algorithm for near-optimal reinforcement learning
Brafman, R. I. and Tennenholtz, M · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, S. M · 2003
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Singh, S. P., Barto, A. G., and Chentanez, N · 2004
Earlier work this paper cites.
Exploration via modelbased interval estimation, 2004
Strehl, A. and Littman, M · 2004
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Bertsekas, D. P · 2005
Earlier work this paper cites.
Prediction, learning, and games
Cesa-Bianchi, N. and Lugosi, G · 2006
Earlier work this paper cites.
Bias and variance approximation in value function estimates
Mannor, S., Simester, D., Sun, P., and Tsitsiklis, J. N · 2007
Cited alongside, same era.
Driven by compression progress: A simple principle explains essential aspects of subjective beauty, novelty, surprise, interestingness, attention, curiosity, creativity, art, science, music, jokes
Schmidhuber, J · 2009
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P · 2010
Cited alongside, same era.
Interval estimation for reinforcement-learning algorithms in continuous-state domains
White, M. and White, A · 2010
Cited alongside, same era.
Mean-variance optimization in Markov decision processes
Mannor, S. and Tsitsiklis, J · 2011
Cited alongside, same era.
From bandits to monte-carlo tree search: The optimistic principle applied to optimization and planning
Munos, R · 2014
Later among the works it cites.
Generalization and exploration via randomized value functions
Osband, I., Van Roy, B., and Wen, Z · 2014
Later among the works it cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Later among the works it cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Later among the works it cites.
Regret analysis of the anytime optimally confident ucb algorithm
Lattimore, T · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the sample complexity of reinforcement learning with a generative model
Azar, M. G., Munos, R., and Kappen, B · 2012
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2012
Cited alongside, same era.
Matrix computations , volume 3
Golub, G. H. and Van Loan, C. F · 2012
Cited alongside, same era.
Efficient Bayes-adaptive reinforcement learning using sample-based search
Guez, A., Silver, D., and Dayan, P · 2012
Cited alongside, same era.
On Bayesian upper confidence bounds for bandit problems
Kaufmann, E., Cappé, O., and Garivier, A · 2012
Cited alongside, same era.
Pac bounds for discounted mdps
Lattimore, T. and Hutter, M · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Cited alongside, same era.
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Later among the works it cites.
On lower bounds for regret in reinforcement learning
Osband, I. and Van Roy, B · 2016
Later among the works it cites.
Deep exploration via bootstrapped DQN
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B · 2016
Later among the works it cites.
Learning the variance of the reward-to-go
Tamar, A., Di Castro, D., and Mannor, S · 2016
Later among the works it cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Closest in time.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Closest in time.
Efficient exploration with double uncertain value networks
Moerland, T. M., Broekens, J., and Jonker, C. M · 2017
Closest in time.
Combining policy gradient and Q-learning
O’Donoghue, B., Munos, R., Kavukcuoglu, K., and Mnih, V · 2017
Closest in time.
Why is posterior sampling better than optimism for reinforcement learning
Osband, I. and Van Roy, B · 2017
Closest in time.
Deep exploration via randomized value functions
Osband, I., Russo, D., Wen, Z., and Van Roy, B · 2017
Closest in time.
Count-based exploration with neural density models
Ostrovski, G., Bellemare, M. G., Oord, A. v. d., and Munos, R · 2017
Closest in time.