Fetching the paper…
Reading the bibliography…
Q-learning with neural network function approximation (neural Q-learning for short) is among the most prevalent deep reinforcement learning algorithms.
A theoretical analysis of deep q-learning
Yang, Z · 1901
Earlier work this paper cites.
A generalization theory of gradient descent for learning over-parameterized deep relu networks
Cao, Y · 1902
Earlier work this paper cites.
Finite-time error bounds for linear stochastic approximation and td learning
Srikant, R · 1902
Earlier work this paper cites.
A gram-gauss-newton method learning overparameterized deep neural networks for regression problems
Cai, T · 1905
Earlier work this paper cites.
Performance of q-learning with linear function approximation: Stability and finite-time analysis
Chen, Z · 1905
Earlier work this paper cites.
Hu, B · 1906
Earlier work this paper cites.
Q-learning
Watkins, C. J · 1992
Earlier work this paper cites.
Convergence of stochastic iterative dynamic programming algorithms
Jaakkola, T · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L · 1995
Earlier work this paper cites.
Dynamic programming and optimal control
Bertsekas, D. P · 1995
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
Tsitsiklis, J. N · 1997
Earlier work this paper cites.
The ode method for convergence of stochastic approximation and reinforcement learning
Borkar, V. S · 2000
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R · 2000
Earlier work this paper cites.
Kernel-based reinforcement learning
Ormoneit, D · 2002
Earlier work this paper cites.
On the existence of fixed points for q-learning and sarsa in partially observable domains
Perkins, T. J · 2002
Cited alongside, same era.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
Riedmiller, M · 2005
Cited alongside, same era.
An analysis of reinforcement learning with function approximation
Melo, F. S · 2008
Cited alongside, same era.
Finite-time bounds for fitted value iteration
Munos, R · 2008
Cited alongside, same era.
Q-learning and pontryagin’s minimum principle
Mehta, P · 2009
Cited alongside, same era.
Algorithms for reinforcement learning
Szepesvari, C · 2010
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Wang, Z · 2016
Later among the works it cites.
Zap q-learning
Devraj, A. M · 2017
Later among the works it cites.
Markov chains and mixing times
Levin, D. A · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
Silver, D · 2017
Later among the works it cites.
A finite time analysis of temporal difference learning with linear function approximation
Bhandari, J · 2018
Later among the works it cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Chizat, L · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Markov chains and stochastic stability
Meyn, S. P · 2012
Cited alongside, same era.
Learning contact-rich manipulation skills with guided policy search
Levine, S · 2015
Cited alongside, same era.
Finite-sample analysis of proximal gradient td algorithms
Liu, B · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V · 2015
Cited alongside, same era.
Deep learning in neural networks: An overview
Schmidhuber, J · 2015
Cited alongside, same era.
Safe, multi-agent, reinforcement learning for autonomous driving
Shalev-Shwartz, S · 2016
Cited alongside, same era.
Dalal, G · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A · 2018
Later among the works it cites.
Scalable deep reinforcement learning for vision-based robotic manipulation
Kalashnikov, D · 2018
Later among the works it cites.
Linear stochastic approximation: How far does constant step-size and iterate averaging go?
Lakshminarayanan, C · 2018
Later among the works it cites.
Planning and decision-making for autonomous vehicles
Schwarting, W · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S · 2018
Later among the works it cites.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Arora, S · 2019
Closest in time.
An improved analysis of training over-parameterized deep neural networks
Zou, D · 2019
Closest in time.