Fetching the paper…
Reading the bibliography…
Exploration remains a key challenge in deep reinforcement learning (RL).
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R · 1933
Earlier work this paper cites.
Exploration bonuses and dual control
Dayan, P. and Sejnowski, T. J · 1996
Earlier work this paper cites.
Reinforcement Learning: an Introduction
Sutton, R. and Barto, A · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 1999
Earlier work this paper cites.
A Bayesian framework for reinforcement learning
Strens, M · 2000
Earlier work this paper cites.
A natural policy gradient
Kakade, S · 2001
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, P · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S · 2002
Earlier work this paper cites.
On actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N · 2003
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Singh, S. P., Barto, A. G., and Chentanez, N · 2004
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Auer, P., Jaksch, T., and Ortner, R · 2008
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Strehl, A. L. and Littman, M. L · 2008
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, B. D · 2010
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2012
Earlier work this paper cites.
Elements of information theory
Cover, T. M. and Thomas, J. A · 2012
Earlier work this paper cites.
Intrinsic motivation and reinforcement learning
Barto, A. G · 2013
Earlier work this paper cites.
(More) efficient reinforcement learning via posterior sampling
Osband, I., Russo, D., and Van Roy, B · 2013
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
Bayesian reinforcement learning: A survey
Ghavamzadeh, M., Mannor, S., Pineau, J., and Tamar, A · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Earlier work this paper cites.
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, B. C., Levine, S., and Abbeel, P · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M · 2016
Cited alongside, same era.
Deep Exploration via Randomized Value Functions
Osband, I · 2016
Cited alongside, same era.
Deep exploration via bootstrapped DQN
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D · 2016
The uncertainty Bellman equation and exploration
O’Donoghue, B., Osband, I., Munos, R., and Mnih, V · 2018
Later among the works it cites.
Randomized prior functions for deep reinforcement learning
Osband, I., Aslanides, J., and Cassirer, A · 2018
Later among the works it cites.
A tutorial on Thompson sampling
Russo, D. J., Van Roy, B., Kazerouni, A., Osband, I., and Wen, Z · 2018
Later among the works it cites.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al · 2019
Later among the works it cites.
Global optimality guarantees for policy gradient methods
Bhandari, J. and Russo, D · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R · 2017
Cited alongside, same era.
Noisy networks for exploration
Fortunato, M., Azar, M. G., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, O., et al · 2017
Cited alongside, same era.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D · 2017
Cited alongside, same era.
A unified view of entropy-regularized markov decision processes
Neu, G., Jonsson, A., and Gómez, V · 2017
Cited alongside, same era.
Combining policy gradient and Q-learning
O’Donoghue, B., Munos, R., Kavukcuoglu, K., and Mnih, V · 2017
Cited alongside, same era.
Count-based exploration with neural density models
Ostrovski, G., Bellemare, M. G., Oord, A. v. d., and Munos, R · 2017
Cited alongside, same era.
Eriksson, H. and Dimitrakakis, C · 2019
Later among the works it cites.
If maxent RL is the answer, what is the question?
Eysenbach, B. and Levine, S · 2019
Later among the works it cites.
Behaviour suite for reinforcement learning
Osband, I., Doron, Y., Hessel, M., Aslanides, J., Sezener, E., Saraiva, A., McKinney, K., Lattimore, T., Szepesvari, C., Singh, S., Roy, B. V., Sutton, R., Silver, D., and Hasselt, H. V · 2019
Later among the works it cites.
Grandmaster level in Starcraft II using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Later among the works it cites.
Temporally-extended ϵ \epsilon -greedy exploration
Dabney, W., Ostrovski, G., and Barreto, A · 2020
Later among the works it cites.
On the global convergence rates of softmax policy gradient methods
Mei, J., Xiao, C., Szepesvari, C., and Schuurmans, D · 2020
Later among the works it cites.
Stochastic matrix games with bandit feedback
O’Donoghue, B., Lattimore, T., and Osband, I · 2020
Later among the works it cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G · 2021
Later among the works it cites.
On the linear convergence of policy gradient methods for finite mdps
Bhandari, J. and Russo, D · 2021
Later among the works it cites.
Podracer architectures for scalable reinforcement learning
Hessel, M., Kroiss, M., Clark, A., Kemaev, I., Quan, J., Keck, T., Viola, F., and van Hasselt, H · 2021
Later among the works it cites.
Reinforcement learning, bit by bit
Lu, X., Van Roy, B., Dwaracherla, V., Ibrahimi, M., Osband, I., and Wen, Z · 2021
Later among the works it cites.
Variational Bayesian reinforcement learning with regret bounds
O’Donoghue, B · 2021
Later among the works it cites.
Variational Bayesian optimistic sampling
O’Donoghue, B. and Lattimore, T · 2021
Later among the works it cites.
The neural testbed: Evaluating joint predictions
Osband, I., Wen, Z., Asghari, S. M., Dwaracherla, V., Hao, B., Ibrahimi, M., Lawson, D., Lu, X., O’Donoghue, B., and Van Roy, B · 2021
Later among the works it cites.
Sample efficient reinforcement learning with REINFORCE
Zhang, J., Kim, J., O’Donoghue, B., and Boyd, S · 2021
Later among the works it cites.
Ensembles for uncertainty estimation: Benefits of prior functions and bootstrapping
Dwaracherla, V., Wen, Z., Osband, I., Lu, X., Asghari, S. M., and Van Roy, B · 2022
Later among the works it cites.
Control Systems and Reinforcement Learning
Meyn, S · 2022
Later among the works it cites.
On the connection between Bregman divergence and value in regularized Markov decision processes
O’Donoghue, B · 2022
Later among the works it cites.