Fetching the paper…
Reading the bibliography…
While there are convergence guarantees for temporal difference (TD) learning when using linear function approximators, the situation for nonlinear models is far less understood, and divergent examples are known.
A theoretical analysis of deep q-learning
Zhuoran Yang, Yuchen Xie, and Zhaoran Wang · 1901
Earlier work this paper cites.
Temporal-difference learning for nonlinear value function approximation in the lazy training regime
Andrea Agazzi and Jianfeng Lu · 1905
Earlier work this paper cites.
Neural temporal-difference learning converges to global optima
Qi Cai, Zhuoran Yang, Jason D. Lee, and Zhaoran Wang · 1905
Earlier work this paper cites.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S. Sutton · 1988
Earlier work this paper cites.
Topics in matrix analysis
Roger A. Horn and Charles R. Johnson · 1991
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
Stochastic approximation with two time scales
Vivek S. Borkar · 1997
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
John N. Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
Performance bounds in l p {}_{\mbox{p}} -norm for approximate value iteration
Rémi Munos · 2007
Earlier work this paper cites.
Stochastic Approximation: A Dynamical Systems Viewpoint
V.S. Borkar · 2008
Cited alongside, same era.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Cited alongside, same era.
Convergent temporal-difference learning with arbitrary smooth function approximation
Shalabh Bhatnagar, Doina Precup, David Silver, Richard S Sutton, Hamid R. Maei, and Csaba Szepesvári · 2009
Cited alongside, same era.
Gradient temporal-difference learning algorithms
Hamid R. Maei · 2011
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Deep residual learning for image recognition
SBEED: Convergent reinforcement learning with nonlinear function approximation
Bo Dai, Albert Shaw, Lihong Li, Lin Xiao, Niao He, Zhen Liu, Jianshu Chen, and Le Song · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Later among the works it cites.
Approximate temporal difference learning is a gradient descent for reversible policies
Yann Ollivier · 2018
Later among the works it cites.
Are resnets provably better than linear predictors?
Ohad Shamir · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Fisher-rao metric, geometry, and complexity of neural networks
Tengyuan Liang, Tomaso A. Poggio, Alexander Rakhlin, and James Stokes · 2017
Cited alongside, same era.
A finite time analysis of temporal difference learning with linear function approximation
Jalaj Bhandari, Daniel Russo, and Raghav Singal · 2018
Cited alongside, same era.
A Note on Lazy Training in Supervised Differentiable Programming
Lénaïc Chizat and Francis Bach · 2018
Cited alongside, same era.
Convergent reinforcement learning with function approximation: A bilevel optimization perspective, 2019a
Zhuoran Yang, Zuyue Fu, Kaiqing Zhang, and Zhaoran Wang
Cited in the paper.
Joshua Achiam, Ethan Knight, and Pieter Abbeel · 2019
Closest in time.
Two-timescale networks for nonlinear value function approximation
Wesley Chung, Somjit Nath, Ajin Joseph, and Martha White · 2019
Closest in time.
Diagnosing Bottlenecks in Deep Q-learning Algorithms
Justin Fu, Aviral Kumar, Matthew Soh, and Sergey Levine · 2019
Closest in time.
Samet Oymak and Mahdi Soltanolkotabi · 2019
Closest in time.