Fetching the paper…
Reading the bibliography…
The question of how to determine which states and actions are responsible for a certain outcome is known as the credit assignment problem and remains a central research question in reinforcement learning and artificial intelligence.
General non-linear Bellman equations
van Hasselt, H. P.; Quan, J.; Hessel, M.; Xu, Z.; Borsa, D.; and Barreto, A. 2019 · 1907
Earlier work this paper cites.
Dynamic Programming
Bellman, R. 1957 · 1957
Earlier work this paper cites.
Steps Toward Artificial Intelligence
Minsky, M. 1963 · 1963
Earlier work this paper cites.
Temporal Credit Assignment in Reinforcement Learning
Sutton, R. S. 1984 · 1984
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S. 1988 · 1988
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, R. S. 1990 · 1990
Earlier work this paper cites.
A menu of designs for reinforcement learning over time
Werbos, P. J. 1990 · 1990
Earlier work this paper cites.
Learning to perceive and act by trial and error
Whitehead, S. D.; and Ballard, D. H. 1991 · 1991
Earlier work this paper cites.
The convergence of TD ( λ ) (\lambda) for general lambda
Dayan, P. 1992 · 1992
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L. 1992 · 1992
Earlier work this paper cites.
Practical Issues in Temporal Difference Learning
Tesauro, G. 1992 · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. 1992 · 1992
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Dayan, P. 1993 · 1993
Earlier work this paper cites.
Prioritized Sweeping: Reinforcement Learning with less Data and less Time
Moore, A. W.; and Atkeson, C. G. 1993 · 1993
Earlier work this paper cites.
Efficient dynamic programming-based learning for control
Peng, J. 1993 · 1993
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L. 1994 · 1994
Earlier work this paper cites.
TD-Gammon, a self-teaching backgammon program, achieves master-level play
Tesauro, G. J. 1994 · 1994
Earlier work this paper cites.
Asynchronous stochastic approximation and Q-learning
Tsitsiklis, J. N. 1994 · 1994
Earlier work this paper cites.
Planning and Acting in Partially Observable Stochastic Domains
Kaelbling, L. P.; Littman, M. L.; and Cassandra, A. R. 1995 · 1995
Earlier work this paper cites.
Neuro-dynamic Programming
Bertsekas, D. P.; and Tsitsiklis, J. N. 1996 · 1996
Earlier work this paper cites.
Incremental Multi-step Q-learning
Peng, J.; and Williams, R. J. 1996 · 1996
Cited alongside, same era.
Reinforcement Learning with replacing eligibility traces
Singh, S. P.; and Sutton, R. S. 1996 · 1996
Cited alongside, same era.
Adaptive critic designs
Prokhorov, D. V.; and Wunsch, D. C. 1997 · 1997
Cited alongside, same era.
An analysis of temporal-difference learning with function approximation
Tsitsiklis, J. N.; and Van Roy, B. 1997 · 1997
Cited alongside, same era.
Natural gradient works efficiently in learning
Amari, S. I. 1998 · 1998
Cited alongside, same era.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S.; McAllester, D.; Singh, S.; and Mansour, Y. 2000 · 2000
Cited alongside, same era.
Prioritized Experience Replay
Schaul, T.; Quan, J.; Antonoglou, I.; and Silver, D. 2016 · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D.; Huang, A.; Maddison, C. J.; Guez, A.; Sifre, L.; Van Den Driessche, G.; Schrittwieser, J.; Antonoglou, I.; Panneershelvam, V.; Lanctot, M.; et al. 2016 · 2016
Later among the works it cites.
An Emphatic Approach to the Problem of Off-policy Temporal-Difference Learning
Sutton, R. S.; Mahmood, A. R.; and White, M. 2016 · 2016
Later among the works it cites.
Learning values across many orders of magnitude
van Hasselt, H. P.; Guez, A.; Hessel, M.; Mnih, V.; and Silver, D. 2016 · 2016
Later among the works it cites.
Dueling Network Architectures for Deep Reinforcement Learning
Wang, Z.; de Freitas, N.; Schaul, T.; Hessel, M.; van Hasselt, H. P.; and Lanctot, M. 2016 · 2016
Later among the works it cites.
Successor features for transfer in reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neural Fitted Q Iteration - First Experiences with a Data Efficient Neural Reinforcement Learning Method
Riedmiller, M. 2005 · 2005
Cited alongside, same era.
Generalized off-policy actor-critic
Zhang, S.; Boehmer, W.; and Whiteson, S. 2019 · 2011
Cited alongside, same era.
Reinforcement Learning in Continuous State and Action Spaces
van Hasselt, H. P. 2012 · 2012
Cited alongside, same era.
The Arcade Learning Environment: An Evaluation Platform for General Agents
Bellemare, M. G.; Naddaf, Y.; Veness, J.; and Bowling, M. 2013 · 2013
Cited alongside, same era.
Planning by Prioritized Sweeping with Small Backups
van Seijen, H.; and Sutton, R. S. 2013 · 2013
Cited alongside, same era.
Off-policy TD( λ \lambda ) with a true online equivalence
van Hasselt, H. P.; Mahmood, A. R.; and Sutton, R. S. 2014 · 2014
Cited alongside, same era.
Barreto, A.; Dabney, W.; Munos, R.; Hunt, J. J.; Schaul, T.; van Hasselt, H. P.; and Silver, D. 2017 · 2017
Later among the works it cites.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
Elfwing, S.; Uchibe, E.; and Doya, K. 2018 · 2018
Later among the works it cites.
Rainbow: Combining Improvements in Deep Reinforcement Learning
Hessel, M.; Modayil, J.; van Hasselt, H. P.; Schaul, T.; Ostrovski, G.; Dabney, W.; Horgan, D.; Piot, B.; Azar, M.; and Silver, D. 2018 · 2018
Later among the works it cites.
Distributed Prioritized Experience Replay
Horgan, D.; Quan, J.; Budden, D.; Barth-Maron, G.; Hessel, M.; van Hasselt, H. P.; and Silver, D. 2018 · 2018
Later among the works it cites.
An Off-policy Policy Gradient Theorem Using Emphatic Weightings
Imani, E.; Graves, E.; and White, M. 2018 · 2018
Later among the works it cites.
Recurrent experience replay in distributed reinforcement learning
Kapturowski, S.; Ostrovski, G.; Quan, J.; Munos, R.; and Dabney, W. 2018 · 2018
Later among the works it cites.
Approximate Temporal Difference Learning is a Gradient Descent for Reversible Policies
Ollivier, Y. 2018 · 2018
Later among the works it cites.
Source Traces for Temporal Difference Learning
Pitis, S. 2018 · 2018
Later among the works it cites.
Observe and look further: Achieving consistent performance on Atari
Pohlen, T.; Piot, B.; Hester, T.; Azar, M. G.; Horgan, D.; Budden, D.; Barth-Maron, G.; van Hasselt, H. P.; Quan, J.; Večerík, M.; Hessel, M.; Munos, R.; and Pietquin, O. 2018 · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S.; and Barto, A. G. 2018 · 2018
Later among the works it cites.
When to use parametric models in reinforcement learning?
van Hasselt, H. P.; Hessel, M.; and Aslanides, J. 2019 · 2019
Later among the works it cites.
Geometric Insights into the Convergence of Non-linear TD Learning
Brandfonbrener, D.; and Bruna, J. 2020 · 2020
Closest in time.
Haiku: Sonnet for JAX
Hennigan, T.; Cai, T.; Norman, T.; and Babuschkin, I. 2020 · 2020
Closest in time.
Deep reinforcement learning with double Q-Learning
van Hasselt, H. P.; Guez, A.; and Silver, D. 2016 · 2094
Closest in time.