Fetching the paper…
Reading the bibliography…
Value function estimation is an important task in reinforcement learning, i.e., prediction.
Sur les opérations dans les ensembles abstraits et leur application aux équations intégrales
Banach, S · 1922
Earlier work this paper cites.
Dynamic programming
Bellman, R. E · 1957
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Learning from delayed rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
Adapting bias by gradient descent: An incremental version of delta-bar-delta
Sutton, R. S · 1992
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
A generalized reinforcement-learning model: Convergence and applications
Littman, M. L. and Szepesvári, C · 1996
Earlier work this paper cites.
Reinforcement learning: A survey
Littman, M. L., Moore, A. W., et al · 1996
Earlier work this paper cites.
Introduction to reinforcement learning , volume 135
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Convergence results for single-step on-policy reinforcement-learning algorithms
Singh, S., Jaakkola, T., Littman, M. L., and Szepesvári, C · 2000
Earlier work this paper cites.
Asymptotic expansions of the hurwitz–lerch zeta function
Ferreira, C. and López, J. L · 2004
Earlier work this paper cites.
Double q-learning
Hasselt, H. V · 2010
Cited alongside, same era.
Dynamic policy programming
Azar, M. G., Gómez, V., and Kappen, H. J · 2012
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Cited alongside, same era.
Reinforcement learning in robotics: A survey
Kober, J., Bagnell, J. A., and Peters, J · 2013
Cited alongside, same era.
van Hasselt, H · 2013
Cited alongside, same era.
Taming the noise in reinforcement learning via soft updates
Fox, R., Pakman, A., and Tishby, N · 2015
Combining policy gradient and q-learning
O’Donoghue, B., Munos, R., Kavukcuoglu, K., and Mnih, V · 2016
Later among the works it cites.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Later among the works it cites.
Boltzmann exploration done right
Cesa-Bianchi, N., Gentile, C., Lugosi, G., and Neu, G · 2017
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., Van Hasselt, H., Lanctot, M., and De Freitas, N · 2015
Cited alongside, same era.
An alternative softmax operator for reinforcement learning
Asadi, K. and Littman, M. L · 2016
Cited alongside, same era.
Increasing the action gap: New operators for reinforcement learning
Bellemare, M. G., Ostrovski, G., Guez, A., Thomas, P. S., and Munos, R · 2016
Cited alongside, same era.
Estimating maximum expected value through gaussian approximation
D’Eramo, C., Restelli, M., and Nuara, A · 2016
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Later among the works it cites.
Sbeed: Convergent reinforcement learning with nonlinear function approximation
Dai, B., Shaw, A., Li, L., Xiao, L., He, N., Liu, Z., Chen, J., and Song, L · 2018
Later among the works it cites.
Learning to walk via deep reinforcement learning
Haarnoja, T., Zhou, A., Ha, S., Tan, J., Tucker, G., and Levine, S · 2018
Later among the works it cites.
Revisiting the softmax bellman operator: Theoretical properties and practical benefits
Song, Z., Parr, R. E., and Carin, L · 2018
Later among the works it cites.
Meta-gradient reinforcement learning
Xu, Z., van Hasselt, H., and Silver, D · 2018
Later among the works it cites.