Fetching the paper…
Reading the bibliography…
Temporal-difference and Q-learning play a key role in deep reinforcement learning, where they are empowered by expressive nonlinear function approximators such as neural networks.
Arora, S · 1901
Earlier work this paper cites.
Analysis of a two-layer neural network via displacement convexity
Javanmard, A · 1901
Earlier work this paper cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J · 1902
Earlier work this paper cites.
Mean-field theory of two-layers neural networks: Dimension-free bounds and kernel limit
Mei, S · 1902
Earlier work this paper cites.
Finite-time error bounds for linear stochastic approximation and TD learning
Srikant, R · 1902
Earlier work this paper cites.
Finite-sample analysis for SARSA and Q-learning with linear function approximation
Zou, S · 1902
Earlier work this paper cites.
Temporal-difference learning for nonlinear value function approximation in the lazy training regime
Agazzi, A · 1905
Earlier work this paper cites.
Geometric insights into the convergence of nonlinear TD learning
Brandfonbrener, D · 1905
Earlier work this paper cites.
On the expected dynamics of nonlinear TD learning
Brandfonbrener, D · 1905
Earlier work this paper cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Cao, Y · 1905
Earlier work this paper cites.
Performance of q-learning with linear function approximation: Stability and finite-time analysis
Chen, Z · 1905
Earlier work this paper cites.
A mean-field limit for certain deep neural networks
Araújo, D · 1906
Earlier work this paper cites.
Ji, Z · 1909
Earlier work this paper cites.
Beyond linearization: On quadratic and higher-order approximation of wide neural networks
Bai, Y · 1910
Earlier work this paper cites.
Over parameterized two-level neural networks can learn near optimal feature representations
Fang, C · 1910
Earlier work this paper cites.
How much over-parameterization is sufficient to learn deep ReLU networks?
Chen, Z · 1911
Earlier work this paper cites.
Convex formulation of overparameterized deep neural networks
Fang, C · 1911
Earlier work this paper cites.
Asymptotics of reinforcement learning with neural networks
Sirignano, J · 1911
Earlier work this paper cites.
Learning distributed representations of concepts
Hinton, G · 1986
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Topics in propagation of chaos
Sznitman, A.-S · 1989
Earlier work this paper cites.
Finite-dimensional variational inequality and nonlinear complementarity problems: A survey of theory, algorithms and applications
Harker, P. T · 1990
Earlier work this paper cites.
Q-learning
Watkins, C · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Barron, A. R · 1993
Cited alongside, same era.
Convergence of stochastic iterative dynamic programming algorithms
Jaakkola, T · 1994
Cited alongside, same era.
Residual algorithms: Reinforcement learning with function approximation
Baird, L · 1995
Cited alongside, same era.
Generalization in reinforcement learning: Safely approximating the value function
Boyan, J. A · 1995
Cited alongside, same era.
Analysis of temporal-diffference learning with function approximation
Tsitsiklis, J. N · 1997
Cited alongside, same era.
Approximation theory of the MLP model in neural networks
Pinkus, A · 1999
Cited alongside, same era.
Deep learning
LeCun, Y · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Mnih, V · 2015
Later among the works it cites.
End-to-end training of deep visuomotor policies
Levine, S · 2016
Later among the works it cites.
Combining policy gradient and Q-learning
O’Donoghue, B · 2016
Later among the works it cites.
SGD learns the conjugate kernel class of the network
Daniely, A · 2017
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The ODE method for convergence of stochastic approximation and reinforcement learning
Borkar, V. S · 2000
Cited alongside, same era.
Actor-critic algorithms
Konda, V. R · 2000
Cited alongside, same era.
Generalization of an inequality by Talagrand and links with the logarithmic Sobolev inequality
Otto, F · 2000
Cited alongside, same era.
Mean-field analysis of two-layer neural networks: Non-asymptotic rates and generalization bounds
Chen, Z · 2002
Cited alongside, same era.
Stochastic approximation and recursive algorithms and applications
Kushner, H · 2003
Cited alongside, same era.
Topics in optimal transportation
Villani, C · 2003
Cited alongside, same era.
Nachum, O · 2017
Later among the works it cites.
Equivalence between policy gradients and soft Q-learning
Schulman, J · 2017
Later among the works it cites.
A finite time analysis of temporal difference learning with linear function approximation
Bhandari, J · 2018
Later among the works it cites.
Finite sample analyses for TD(0) with function approximation
Dalal, G · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A · 2018
Later among the works it cites.
Linear stochastic approximation: How far does constant step-size and iterate averaging go?
Lakshminarayanan, C · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y · 2018
Later among the works it cites.
A mean field view of the landscape of two-layer neural networks
Mei, S · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep ReLU networks
Zou, D · 2018
Later among the works it cites.
Feature-based aggregation and deep reinforcement learning: A survey and some new implementations
Bertsekas, D. P · 2019
Later among the works it cites.
Neural temporal-difference learning converges to global optima
Cai, Q · 2019
Later among the works it cites.
A course in functional analysis
Conway, J. B · 2019
Later among the works it cites.
High-dimensional statistics: A non-asymptotic viewpoint
Wainwright, M. J · 2019
Later among the works it cites.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
Wei, C · 2019
Later among the works it cites.
Two time-scale off-policy TD learning: Non-asymptotic analysis over Markovian samples
Xu, T · 2019
Later among the works it cites.
An improved analysis of training over-parameterized deep neural networks
Zou, D · 2019
Later among the works it cites.