Fetching the paper…
Reading the bibliography…
In value-based reinforcement learning (RL), unlike in supervised learning, the agent faces not a single, stationary, approximation problem, but a sequence of value prediction problems.
Dynamic Programming and Markov Processes
Howard, R. 1960 · 1960
Earlier work this paper cites.
Q-learning
Watkins, C. J.; and Dayan, P. 1992 · 1992
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Dayan, P. 1993 · 1993
Earlier work this paper cites.
Markov Decision Processes
Puterman, M. L. 1994 · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L. 1995 · 1995
Earlier work this paper cites.
Neuro-dynamic programming , volume 5
Bertsekas, D. P.; and Tsitsiklis, J. N. 1996 · 1996
Earlier work this paper cites.
Reinforcement Learning with Selective Perception and Hidden State
McCallum, A. K. 1996 · 1996
Earlier work this paper cites.
Error bounds for approximate policy iteration
Munos, R. 2003 · 2003
Earlier work this paper cites.
Efficient algorithms for online decision problems
Kalai, A.; and Vempala, S. 2005 · 2005
Earlier work this paper cites.
Proto-value functions: Developmental reinforcement learning
Mahadevan, S. 2005 · 2005
Earlier work this paper cites.
Towards a unified theory of state abstraction for MDPs
Li, L.; Walsh, T.; and Littman, M. 2006 · 2006
Earlier work this paper cites.
An analysis of linear models, linear value-function approximation, and feature selection for reinforcement learning
Parr, R.; Li, L.; Taylor, G.; Painter-Wakefield, C.; and Littman, M. L. 2008 · 2008
Earlier work this paper cites.
Regularization and feature selection in least-squares temporal difference learning
Kolter, J. Z.; and Ng, A. Y. 2009 · 2009
Cited alongside, same era.
Value function approximation in reinforcement learning using the Fourier basis
Konidaris, G.; Osentoski, S.; and Thomas, P. 2011 · 2011
Cited alongside, same era.
On the relation of slow feature analysis and Laplacian eigenmaps
Sprekeler, H. 2011 · 2011
Cited alongside, same era.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S.; Modayil, J.; Delp, M.; Degris, T.; Pilarski, P. M.; White, A.; and Precup, D. 2011 · 2011
Cited alongside, same era.
The simplex and policy-iteration methods are strongly polynomial for the Markov decision problem with a fixed discount rate
Ye, Y. 2011 · 2011
Cited alongside, same era.
A distributional perspective on reinforcement learning
Bellemare, M. G.; Dabney, W.; and Munos, R. 2017 · 2017
Later among the works it cites.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M.; Mnih, V.; Czarnecki, W. M.; Schaul, T.; Leibo, J. Z.; Silver, D.; and Kavukcuoglu, K. 2017 · 2017
Later among the works it cites.
Distributed Distributional Deterministic Policy Gradients
Barth-Maron, G.; Hoffman, M. W.; Budden, D.; Dabney, W.; Horgan, D.; TB, D.; Muldal, A.; Heess, N.; and Lillicrap, T. 2018 · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M.; Modayil, J.; Van Hasselt, H.; Schaul, T.; Ostrovski, G.; Dabney, W.; Horgan, D.; Piot, B.; Azar, M.; and Silver, D. 2018 · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S.; and Barto, A. G. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bellemare, M. G.; Naddaf, Y.; Veness, J.; and Bowling, M. 2013 · 2013
Cited alongside, same era.
Optimization, learning, and games with predictable sequences
Rakhlin, S.; and Sridharan, K. 2013 · 2013
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M.; Fidjeland, A. K.; Ostrovski, G.; et al. 2015 · 2015
Cited alongside, same era.
Universal value function approximators
Schaul, T.; Horgan, D.; Gregor, K.; and Silver, D. 2015 · 2015
Cited alongside, same era.
Linear feature encoding for reinforcement learning
Song, Z.; Parr, R. E.; Liao, X.; and Carin, L. 2016 · 2016
Cited alongside, same era.
Deep reinforcement learning with double Q-learning
van Hasselt, H.; Guez, A.; and Silver, D. 2016 · 2016
Cited alongside, same era.
Successor Features for Transfer in Reinforcement Learning
Barreto, A.; Dabney, W.; Munos, R.; Hunt, J.; Schaul, T.; van Hasselt, H.; and Silver, D. 2017 · 2017
Cited alongside, same era.
Van Hasselt, H.; Doron, Y.; Strub, F.; Hessel, M.; Sonnerat, N.; and Modayil, J. 2018 · 2018
Later among the works it cites.
A Geometric Perspective on Optimal Representations for Reinforcement Learning
Bellemare, M. G.; Dabney, W.; Dadashi, R.; Taiga, A. A.; Castro, P. S.; Roux, N. L.; Schuurmans, D.; Lattimore, T.; and Lyle, C. 2019 · 2019
Later among the works it cites.
Universal Successor Features Approximators
Borsa, D.; Barreto, A.; Quan, J.; Mankowitz, D. J.; van Hasselt, H.; Munos, R.; Silver, D.; and Schaul, T. 2019 · 2019
Later among the works it cites.
The value function polytope in reinforcement learning
Dadashi, R.; Taïga, A. A.; Roux, N. L.; Schuurmans, D.; and Bellemare, M. G. 2019 · 2019
Later among the works it cites.
Hyperbolic discounting and learning over multiple horizons
Fedus, W.; Gelada, C.; Bengio, Y.; Bellemare, M. G.; and Larochelle, H. 2019 · 2019
Later among the works it cites.
Statistics and samples in distributional reinforcement learning
Rowland, M.; Dadashi, R.; Kumar, S.; Munos, R.; Bellemare, M. G.; and Dabney, W. 2019 · 2019
Later among the works it cites.