Fetching the paper…
Reading the bibliography…
Delusional bias is a fundamental source of error in approximate Q-learning.
Learning from Delayed Rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
Stable function approximation in dynamic programming
Gordon, G. J · 1995
Earlier work this paper cites.
Statistical Learning Theory
Vapnik, V. N · 1998
Earlier work this paper cites.
Approximation Solutions to Markov Decision Problems
Gordon, G · 1999
Earlier work this paper cites.
Interpolation-based Q-learning
Szepesvári, C. and Smart, W · 2004
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., Geurts, P., and Wehenkel, L · 2005
Earlier work this paper cites.
Neural fitted q iteration—first experiences with a data efficient neural reinforcement learning method
Riedmiller, M · 2005
Earlier work this paper cites.
Q-learning with linear function approximation
Melo, F. and Ribeiro, M. I · 2007
Earlier work this paper cites.
Toward off-policy learning control wtih function approximation
Maei, H., Szepesvári, C., Bhatnagar, S., and Sutton, R · 2010
Earlier work this paper cites.
Double q-learning
van Hasselt, H · 2010
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Neural network ensembles in reinforcement learning
Faußer, S. and Schwenker, F · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A., Veness, J., Bellemare, M., Graves, A., Riedmiller, M., Fidjeland, A., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Hasselt, H. v., Guez, A., and Silver, D · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M · 2016
Cited alongside, same era.
Deep exploration via bootstrapped dqn
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Dopamine: A research framework for deep reinforcement learning
Castro, P. S., Moitra, S., Gelada, C., Kumar, S., and Bellemare, M. G · 2018
Later among the works it cites.
Conti, E., Madhavan, V., Such, F. P., Lehman, J., Stanley, K. O., and Clune, J · 2018
Later among the works it cites.
TF-Agents: A library for reinforcement learning in tensorflow
Guadarrama, S., Korattikara, A., Oscar Ramirez, P. C., Holly, E., Fishman, S., Wang, K., Gonina, E., Wu, N., Harris, C., Vanhoucke, V., and Brevdo, E · 2018
Later among the works it cites.
Evolution-guided policy gradient in reinforcement learning
Khadka, S. and Tumer, K · 2018
Later among the works it cites.
Non-delusional Q-learning and value iteration
Lu, T., Schuurmans, D., and Boutilier, C · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., van Hasselt, H., Lanctot, M., and de Freitas, N · 2016
Cited alongside, same era.
Averaged-dqn: Variance reduction and stabilization for deep reinforcement learning
Anschel, O., Baram, N., and Shimkin, N · 2017
Cited alongside, same era.
A distributional perspective on reinforcement learning
Bellemare, M., Dabney, W., and Munos, R · 2017
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D · 2017
Cited alongside, same era.
Observe and look further: Achieving consistent performance on atari
Pohlen, T., Piot, B., Hester, T., Azar, M. G., Horgan, D., Budden, D., Barth-Maron, G., van Hasselt, H., Quan, J., Vecerík, M., Hessel, M., Munos, R., and Pietquin, O · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
Watkins, C. J. C. H. and Dayan, P · 2018
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., Lillicrap, T., and Silver, D · 2019
Later among the works it cites.
Recurrent experience replay in distributed reinforcement learning
Kapturowski, S., Ostrovski, G., Quan, J., Munos, R., and Dabney, W · 2020
Closest in time.