Fetching the paper…
Reading the bibliography…
Temporal-difference learning (TD), coupled with neural networks, is among the most fundamental building blocks of deep reinforcement learning.
Arora, S · 1901
Earlier work this paper cites.
A generalization theory of gradient descent for learning over-parameterized deep ReLU networks
Cao, Y · 1902
Earlier work this paper cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J · 1902
Earlier work this paper cites.
Finite-time error bounds for linear stochastic approximation and TD learning
Srikant, R · 1902
Earlier work this paper cites.
Finite-sample analysis for SARSA and Q-learning with linear function approximation
Zou, S · 1902
Earlier work this paper cites.
Towards characterizing divergence in deep Q-learning
Achiam, J · 1903
Earlier work this paper cites.
A selective overview of deep learning
Fan, J · 1904
Earlier work this paper cites.
Temporal-difference learning for nonlinear value function approximation in the lazy training regime
Agazzi, A · 1905
Earlier work this paper cites.
Geometric insights into the convergence of nonlinear TD learning
Brandfonbrener, D · 1905
Earlier work this paper cites.
On the expected dynamics of nonlinear TD learning
Brandfonbrener, D · 1905
Earlier work this paper cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Cao, Y · 1905
Earlier work this paper cites.
Convergence of adversarial training in overparametrized networks
Gao, R · 1906
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Finite-dimensional variational inequality and nonlinear complementarity problems: a survey of theory, algorithms and applications
Harker, P. T · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Convergence of stochastic iterative dynamic programming algorithms
Jaakkola, T · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L · 1995
Earlier work this paper cites.
Generalization in reinforcement learning: Safely approximating the value function
Boyan, J. A · 1995
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Bradtke, S. J · 1996
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
Tsitsiklis, J. N · 1997
Earlier work this paper cites.
Least-squares temporal difference learning
Boyan, J. A · 1999
Earlier work this paper cites.
The ODE method for convergence of stochastic approximation and reinforcement learning
Borkar, V. S · 2000
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R · 2000
Cited alongside, same era.
Stochastic Approximation and Recursive Algorithms and Applications
Kushner, H · 2003
Cited alongside, same era.
Finite-Dimensional Variational Inequalities and Complementarity Problems
Facchinei, F · 2007
Cited alongside, same era.
Kernel methods in machine learning
Hofmann, T · 2008
Cited alongside, same era.
An analysis of reinforcement learning with function approximation
Melo, F. S · 2008
Cited alongside, same era.
Convergent temporal-difference learning with arbitrary smooth function approximation
Bhatnagar, S · 2009
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Haarnoja, T · 2017
Later among the works it cites.
Non-convex optimization for machine learning
Jain, P · 2017
Later among the works it cites.
A unified view of entropy-regularized Markov decision processes
Neu, G · 2017
Later among the works it cites.
Equivalence between policy gradients and soft Q-learning
Schulman, J · 2017
Later among the works it cites.
Convergent tree-backup and retrace with function approximation
Touati, A · 2017
Later among the works it cites.
Least-squares temporal difference learning for the linear quadratic regulator
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Borkar, V. S · 2009
Cited alongside, same era.
LSTD with random projections
Ghavamzadeh, M · 2010
Cited alongside, same era.
Finite-sample analysis of LSTD
Lazaric, A · 2010
Cited alongside, same era.
Algorithms for reinforcement learning
Szepesvári, C · 2010
Cited alongside, same era.
Algorithmic survey of parametric value function approximation
Geist, M · 2013
Cited alongside, same era.
Policy evaluation with temporal differences: A survey and comparison
Dann, C · 2014
Cited alongside, same era.
Tu, S · 2017
Later among the works it cites.
Finite sample analysis of the GTD policy evaluation algorithms in Markov setting
Wang, Y · 2017
Later among the works it cites.
TD or not TD: Analyzing the role of temporal differencing in deep reinforcement learning
Amiranashvili, A · 2018
Later among the works it cites.
A finite time analysis of temporal difference learning with linear function approximation
Bhandari, J · 2018
Later among the works it cites.
A note on lazy training in supervised differentiable programming
Chizat, L · 2018
Later among the works it cites.
Finite sample analyses for TD(0) with function approximation
Dalal, G · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T · 2018
Later among the works it cites.
Deep reinforcement learning that matters
Henderson, P · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A · 2018
Later among the works it cites.
Linear stochastic approximation: How far does constant step-size and iterate averaging go?
Lakshminarayanan, C · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y · 2018
Later among the works it cites.
Towards understanding the role of over-parametrization in generalization of neural networks
Neyshabur, B · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep ReLU networks
Zou, D · 2018
Later among the works it cites.
Feature-based aggregation and deep reinforcement learning: A survey and some new implementations
Bertsekas, D. P · 2019
Closest in time.
Two-timescale networks for nonlinear value function approximation
Chung, W · 2019
Closest in time.