Fetching the paper…
Reading the bibliography…
Reinforcement learning with function approximation can be unstable and even divergent, especially when combined with off-policy learning and Bellman updates.
Adaptive Algorithms and Stochastic Approximations
Benveniste, A., Priouret, P., and Métivier, M · 1990
Earlier work this paper cites.
Temporal-difference methods and markov models
Barnard, E · 1993
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Dayan, P · 1993
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L. C · 1995
Earlier work this paper cites.
Stable function approximation in dynamic programming
Gordon, G. J · 1995
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Bradtke, S. J. and Barto, A. G · 1996
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
Tsitsiklis, J. N. and Roy, B. V · 1996
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S. P · 1999
Earlier work this paper cites.
The o.d.e. method for convergence of stochastic approximation and reinforcement learning
Borkar, V. S. and Meyn, S. P · 2000
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, M. G. and Parr, R · 2003
Earlier work this paper cites.
Spectra and pseudospectra : the behavior of nonnormal matrices and operators
Trefethen, L. N. and Embree, M · 2005
Earlier work this paper cites.
Proto-value functions: A laplacian framework for learning representation and control in markov decision processes
Mahadevan, S. and Maggioni, M · 2007
Earlier work this paper cites.
Analyzing feature generation for value-function approximation
Parr, R., Painter-Wakefield, C., Li, L., and Littman, M. L · 2007
Earlier work this paper cites.
An analysis of laplacian methods for value function approximation in mdps
Petrik, M · 2007
Earlier work this paper cites.
An analysis of linear models, linear value-function approximation, and feature selection for reinforcement learning
Parr, R., Li, L., Taylor, G., Painter-Wakefield, C., and Littman, M. L · 2008
Cited alongside, same era.
Linear system theory: the state space approach
Zadeh, L. and Desoer, C · 2008
Cited alongside, same era.
Convergent temporal-difference learning with arbitrary smooth function approximation
Maei, H. R., Szepesvari, C., Bhatnagar, S., Precup, D., Silver, D., and Sutton, R. S · 2009
Cited alongside, same era.
Basis function adaptation methods for cost approximation in mdp
Yu, H. and Bertsekas, D. P · 2009
Cited alongside, same era.
Approximate policy iteration: a survey and some new methods
Bertsekas, D. P · 2011
Cited alongside, same era.
Feature-based aggregation and deep reinforcement learning: a survey and some new implementations
Bertsekas, D. P · 2018
Later among the works it cites.
Combined reinforcement learning via abstract representations
François-Lavet, V., Bengio, Y., Precup, D., and Pineau, J · 2018
Later among the works it cites.
Eigenoption discovery through the deep successor representation
Machado, M. C., Rosenbaum, C., Guo, X., Liu, M., Tesauro, G., and Campbell, M · 2018
Later among the works it cites.
Approximate temporal difference learning is a gradient descent for reversible policies
Ollivier, Y · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shi, X. and Wei, Y · 2012
Cited alongside, same era.
Matrix Computations
Golub, G. H. and van Loan, C. F · 2013
Cited alongside, same era.
Policy evaluation with temporal differences: a survey and comparison
Dann, C., Neumann, G., and Peters, J · 2014
Cited alongside, same era.
Design principles of the hippocampal cognitive map
Stachenfeld, K. L., Botvinick, M., and Gershman, S. J · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Cited alongside, same era.
van Hasselt, H., Doron, Y., Strub, F., Hessel, M., Sonnerat, N., and Modayil, J · 2018
Later among the works it cites.
The laplacian in rl: Learning representations with efficient approximations
Wu, Y., Tucker, G., and Nachum, O · 2018
Later among the works it cites.
Fast feature selection for linear value function approximation
Behzadian, B., Gharatappeh, S., and Petrik, M · 2019
Later among the works it cites.
A geometric perspective on optimal representations for reinforcement learning
Bellemare, M. G., Dabney, W., Dadashi, R., Taïga, A. A., Castro, P. S., Roux, N. L., Schuurmans, D., Lattimore, T., and Lyle, C · 2019
Later among the works it cites.
Quantile QT-opt for risk-aware vision-based robotic grasping
Bodnar, C., Li, A., Hausman, K., Pastor, P., and Kalakrishnan, M · 2019
Later among the works it cites.
A framework for data-driven robotics
Cabi, S., Colmenarejo, S. G., Novikov, A., Konyushkova, K., Reed, S., Jeong, R., Zolna, K., Aytar, Y., Budden, D., Vecerik, M., Sushkov, O., Barker, D., Scholz, J., Denil, M., de Freitas, N., and Wang, Z · 2019
Later among the works it cites.
Two-timescale networks for nonlinear value function approximation
Chung, W., Nath, S., Joseph, A. G., and White, M · 2019
Later among the works it cites.
DeepMDP: Learning continuous latent space models for representation learning
Gelada, C., Kumar, S., Buckman, J., Nachum, O., and Bellemare, M. G · 2019
Later among the works it cites.
A practical approach to insertion with variable socket position using deep reinforcement learning
Vecerik, M., Sushkov, O., Barker, D., Rothörl, T., Hester, T., and Scholz, J · 2019
Later among the works it cites.