Fetching the paper…
Reading the bibliography…
Exploration remains a fundamental challenge in reinforcement learning, as many existing methods either lack theoretical guarantees or fall short in practical effectiveness.
Adjustment of an inverse matrix corresponding to a change in one element of a given matrix
Sherman, J. and Morrison, J. W · 1950
Earlier work this paper cites.
Note on a method for calculating corrected sums of squares and products
Welford, B. P · 1962
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S · 2002
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Li, L., Chu, W., Langford, J., and Schapire, R. E · 2010
Earlier work this paper cites.
Contextual bandits with linear payoff functions
Chu, W., Li, L., Reyzin, L., and Schapire, R. E · 2011
Earlier work this paper cites.
Thompson sampling for contextual bandits with linear payoffs
Agrawal, S. and Goyal, N · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Finite-time analysis of kernelised contextual bandits
Valko, M., Korda, N., Munos, R., Flaounas, I., and Cristianini, N · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra1, D., Legg, S., and Hassabis, D · 2015
Earlier work this paper cites.
Efficient learning in large-scale combinatorial semi-bandits
Wen, Z., Kveton, B., and Ashkan, A · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M. G., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Earlier work this paper cites.
Generalization and exploration via randomized value functions
Osband, I., Roy, V. B., and Wen, Z · 2016
Earlier work this paper cites.
On kernelized multi-armed bandits
Chowdhury, S. R. and Gopalan, A · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., Chen, Y., and Lillicrap, T. e. a · 2017
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., Legg, S., and Kavukcuoglu, K · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H. v., and Meger, D · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Cited alongside, same era.
On function approximation in reinforcement learning: Optimism in the face of large state spaces
Yang, Z., Jin, C., Wang, Z., Wang, M., and Jordan, M. I · 2020
Later among the works it cites.
Neural contextual bandits with ucb-based exploration
Zhou, D., Li, L., and Gu, Q · 2020
Later among the works it cites.
Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors
Duan, J., Guan, Y., Li, S., Ren, Y., Sun, Q., and Cheng, B · 2021
Later among the works it cites.
Adversarially guided actor-critic
Flet-Berliac, Y., Ferret, J., Pietquin, O., Preux, P., and Geist, M · 2021
Later among the works it cites.
Randomized exploration for reinforcement learning with general value function approximation
Ishfaq, H., Cui, Q., Nguyen, V., Ayoub, A., Zhuoran, Y., Wang, Z., Precup, D., and Yang, F. L · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I · 2018
Cited alongside, same era.
Deep bayesian bandits showdown: An empirical comparison of bayesian deep networks for thompson sampling
Riquelme, C., Tucker, G., and Snoek, J · 2018
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R., and Wang, R · 2019
Cited alongside, same era.
Deep exploration via randomized value functions
Osband, I., Roy, V. B., Russo, J. D., and Wen, Z · 2019
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., and Georgiev, P. e. a · 2019
Cited alongside, same era.
Neural linear bandits: Overcoming catastrophic forgetting through likelihood matching
Zahavy, T. and Mannor, S · 2019
Cited alongside, same era.
Pc-pg: Policy cover directed exploration for provable policy gradient learning
Agarwal, A., Kakade, S., Henaff, M., and Sun, W · 2020
Cited alongside, same era.
Minihack the planet: A sandbox for open-ended reinforcement learning research
Samvelyan, M., Kirk, R., Kurin, V., Parker-Holder, J., Jiang, M., Hambro, E., Petroni, F., Kuttler, H., Grefenstette, E., and Rocktäschel, T · 2021
Later among the works it cites.
Anti-concentrated confidence bonuses for scalable exploration
Ash, J. T., Zhang, C., Goel, S., Krishnamurthy, A., and Kakade, S · 2022
Later among the works it cites.
Exploration via elliptical episodic bonuses
Henaff, M., Raileanu, R., Jiang, M., and Rocktäschel, T · 2022
Later among the works it cites.
Cleanrl: High-quality single-file implementations of deep reinforcement learning algorithms
Huang, S., Dossa, R. F. J., Ye, C., Braga, J., Chakraborty, D., Mehta, K., and Araújo, J. G · 2022
Later among the works it cites.
Neural contextual bandits with deep representation and shallow exploration
Xu, P., Wen, Z., Zhao, H., and Gu, Q · 2022
Later among the works it cites.
Dsac-t: Distributional soft actor-critic with three refinements
Duan, J., Wang, W., Xiao, L., Gao, J., and Li, S · 2023
Later among the works it cites.
A study of global and episodic bonuses for exploration in contextual mdps
Henaff, M., Jiang, M., and Raileanu, R · 2023
Later among the works it cites.
Curiosity in hindsight: Intrinsic exploration in stochastic environments
Jarrett, D., Tallec, C., Altché, F., Mesnard, T., Munos, R., and Valko, M · 2023
Later among the works it cites.
Automatic intrinsic reward shaping for exploration in deep reinforcement learning
Yuan, M., Li, B., Jin, X., and Zeng, W · 2023
Later among the works it cites.
Deep reinforcement learning without experience replay, target networks, or batch updates
Elsayed, M., Vasan, G., and Mahmood, A. R · 2024
Later among the works it cites.
A contextual combinatorial bandit approach to negotiation
Li, Y., Mu, Z., and Qi, S · 2024
Later among the works it cites.