Fetching the paper…
Reading the bibliography…
One principled approach for provably efficient exploration is incorporating the upper confidence bound (UCB) into the value function as a bonus.
Outlier models and prior distributions in bayesian linear regression
West, M · 1984
Earlier work this paper cites.
Estimating the mean and variance of the target probability distribution
Nix, D. A. and Weigend, A. S · 1994
Earlier work this paper cites.
The elements of statistical learning , volume 1
Friedman, J., Hastie, T., and Tibshirani, R · 2001
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
Auer, P. and Ortner, R · 2007
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C · 2011
Earlier work this paper cites.
Sample complexity of episodic fixed-horizon reinforcement learning
Dann, C. and Brunskill, E · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M. A., Fidjeland, A., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Earlier work this paper cites.
Vime: Variational information maximizing exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., De Turck, F., and Abbeel, P · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped dqn
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B · 2016
Earlier work this paper cites.
Posterior sampling for reinforcement learning: worst-case regret bounds
Agrawal, S. and Jia, R · 2017
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R · 2017
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Earlier work this paper cites.
Ucb exploration via q-ensembles
Chen, R. Y., Sidor, S., Abbeel, P., and Schulman, J · 2017
Earlier work this paper cites.
Uncertainty-driven imagination for continuous deep reinforcement learning
Kalweit, G. and Boedecker, J · 2017
Earlier work this paper cites.
Why is posterior sampling better than optimism for reinforcement learning?
Osband, I. and Van Roy, B · 2017
Earlier work this paper cites.
Count-based exploration with neural density models
Ostrovski, G., Bellemare, M. G., van den Oord, A., and Munos, R · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Cited alongside, same era.
Exploration: A study of count-based exploration for deep reinforcement learning
Tang, H., Houthooft, R., Foote, D., Stooke, A., Chen, O. X., Duan, Y., Schulman, J., DeTurck, F., and Abbeel, P · 2017
Cited alongside, same era.
Efficient exploration through bayesian deep q-networks
Azizzadenesheli, K., Brunskill, E., and Anandkumar, A · 2018
Cited alongside, same era.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Buckman, J., Hafner, D., Tucker, G., Brevdo, E., and Lee, H · 2018
Cited alongside, same era.
Noisy networks for exploration
Fortunato, M., Azar, M. G., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, O., Blundell, C., and Legg, S · 2018
Worst-case regret bounds for exploration via randomized value functions
Russo, D · 2019
Later among the works it cites.
Episodic curiosity through reachability
Savinov, N., Raichuk, A., Marinier, R., Vincent, D., Pollefeys, M., Lillicrap, T., and Gelly, S · 2019
Later among the works it cites.
Optimism in reinforcement learning with generalized linear function approximation
Wang, Y., Wang, R., Du, S. S., and Krishnamurthy, A · 2019
Later among the works it cites.
Never give up: Learning directed exploration strategies
Badia, A. P., Sprechmann, P., Vitvitskyi, A., Guo, D., Piot, B., Kapturowski, S., Tieleman, O., Arjovsky, M., Pritzel, A., Bolt, A., and Blundell, C · 2020
Later among the works it cites.
Variational dynamic for self-supervised exploration in deep reinforcement learning
Bai, C., Liu, P., Wang, Z., Liu, K., Wang, L., and Zhao, Y · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D · 2018
Cited alongside, same era.
Is q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I · 2018
Cited alongside, same era.
Randomized prior functions for deep reinforcement learning
Osband, I., Aslanides, J., and Cassirer, A · 2018
Cited alongside, same era.
The uncertainty bellman equation and exploration
O’Donoghue, B., Osband, I., Munos, R., and Mnih, V · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Arora, S., Du, S., Hu, W., Li, Z., and Wang, R · 2019
Cited alongside, same era.
Ready policy one: World building through active learning
Ball, P., Parker-Holder, J., Pacchiano, A., Choromanski, K., and Roberts, S · 2020
Later among the works it cites.
Provably efficient exploration in policy optimization
Cai, Q., Yang, Z., Jin, C., and Wang, Z · 2020
Later among the works it cites.
Efficient model-based reinforcement learning through optimistic policy search and planning
Curi, S., Berkenkamp, F., and Krause, A · 2020
Later among the works it cites.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I · 2020
Later among the works it cites.
Sunrise: A simple unified framework for ensemble learning in deep reinforcement learning
Lee, K., Laskin, M., Srinivas, A., and Abbeel, P · 2020
Later among the works it cites.
On optimism in model-based reinforcement learning
Pacchiano, A., Ball, P., Parker-Holder, J., Choromanski, K., and Roberts, S · 2020
Later among the works it cites.
Optimistic exploration even with a pessimistic initialisation
Rashid, T., Peng, B., Boehmer, W., and Whiteson, S · 2020
Later among the works it cites.
Implicit generative modeling for efficient exploration
Ratzlaff, N., Bai, Q., Fuxin, L., and Xu, W · 2020
Later among the works it cites.
Planning to explore via self-supervised world models
Sekar, R., Rybkin, O., Daniilidis, K., Abbeel, P., Hafner, D., and Pathak, D · 2020
Later among the works it cites.
Optimistic policy optimization with bandit feedback
Shani, L., Efroni, Y., Rosenberg, A., and Mannor, S · 2020
Later among the works it cites.
On bonus based exploration methods in the arcade learning environment
Taiga, A. A., Fedus, W., Machado, M. C., Courville, A., and Bellemare, M. G · 2020
Later among the works it cites.
Frequentist regret bounds for randomized least-squares value iteration
Zanette, A., Brandfonbrener, D., Brunskill, E., Pirotta, M., and Lazaric, A · 2020
Later among the works it cites.