Fetching the paper…
Reading the bibliography…
Discounted reinforcement learning is fundamentally incompatible with function approximation for control in continuing tasks.
Hyperbolic discounting and learning over multiple horizons
Fedus, W., Gelada, C., Bengio, Y., Bellemare, M. G., and Larochelle, H. (2019) · 1902
Earlier work this paper cites.
Issues concerning realizability of Blackwell optimal policies in reinforcement learning
Denis, N. (2019) · 1905
Earlier work this paper cites.
Function approximations and dynamic programming
Bellman, R. and Dreyfus, S. (1959) · 1959
Earlier work this paper cites.
A reinforcement learning method for maximizing undiscounted rewards
Schwartz, A. (1993) · 1993
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L. (1994) · 1994
Earlier work this paper cites.
Learning without state-estimation in partially observable Markovian decision processes
Singh, S. P., Jaakkola, T. S., and Jordan, M. I. (1994) · 1994
Cited alongside, same era.
Average reward reinforcement learning: Foundations, algorithms, and empirical results
Mahadevan, S. (1996) · 1996
Cited alongside, same era.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N. (1999) · 1999
Cited alongside, same era.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (2000) · 2000
Cited alongside, same era.
Learning algorithms for Markov decision processes with average cost
Abounadi, J., Bertsekas, D. P., and Borkar, V. S. (2001) · 2001
Cited alongside, same era.
Dynamic Programming and Optimal Control , vol. II
Bertsekas, D. P. (2005) · 2005
Later among the works it cites.
Examples concerning Abel and Cesàro limits
Bishop, C. J., Feinberg, E. A., and Zhang, J. (2014) · 2014
Later among the works it cites.
Unifying task specification in reinforcement learning
White, M. (2017) · 2017
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…