Fetching the paper…
Reading the bibliography…
Temporal-Difference (TD) learning with nonlinear smooth function approximation for policy evaluation has achieved great success in modern reinforcement learning.
Lu, S · 1902
Earlier work this paper cites.
Finite-time error bounds for linear stochastic approximation and td learning
Srikant, R · 1902
Earlier work this paper cites.
Momentum-based variance reduction in non-convex sgd
Cutkosky, A · 1905
Earlier work this paper cites.
On gradient descent ascent for nonconvex-concave minimax problems
Lin, T · 1906
Earlier work this paper cites.
A multistep lyapunov approach for finite-time analysis of biased stochastic approximation
Wang, G · 1909
Earlier work this paper cites.
A finite-time analysis of q-learning with neural network function approximation
Xu, P · 1912
Earlier work this paper cites.
A linearization method for nonsmooth stochastic programming problems
Ruszczyński, A · 1987
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
Tsitsiklis, J. N · 1997
Earlier work this paper cites.
Luo, L · 2001
Earlier work this paper cites.
Reanalysis of variance reduced temporal difference learning
Xu, T · 2001
Earlier work this paper cites.
Convergent temporal-difference learning with arbitrary smooth function approximation
Bhatnagar, S · 2009
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R · 2013
Cited alongside, same era.
Policy evaluation with temporal differences: A survey and comparison
Dann, C · 2014
Cited alongside, same era.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Defazio, A · 2014
Cited alongside, same era.
Finite-sample analysis of proximal gradient td algorithms
Liu, B · 2015
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
Schulman, J · 2015
Cited alongside, same era.
A finite time analysis of temporal difference learning with linear function approximation
Bhandari, J · 2018
Later among the works it cites.
Two-timescale networks for nonlinear value function approximation
Chung, W · 2018
Later among the works it cites.
A single time-scale stochastic approximation method for nested stochastic optimization
Ghadimi, S · 2018
Later among the works it cites.
Lectures on convex optimization
Nesterov, Y · 2018
Later among the works it cites.
Non-convex min-max optimization: Provable algorithms and applications in machine learning
Rafique, H · 2018
Later among the works it cites.
Neural temporal-difference learning converges to global optima
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dalal, G · 2017
Cited alongside, same era.
Stochastic variance reduction methods for policy evaluation
Du, S. S · 2017
Cited alongside, same era.
Mastering the game of Go without human knowledge
Silver, D · 2017
Cited alongside, same era.
Convergent tree backup and retrace with function approximation
Touati, A · 2017
Cited alongside, same era.
Finite sample analysis of the gtd policy evaluation algorithms in markov setting
Wang, Y · 2017
Cited alongside, same era.
On convergence of some gradient-based temporal-differences algorithms for off-policy learning
Yu, H · 2017
Cited alongside, same era.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Sutton, R. S
Cited in the paper.
Cai, Q · 2019
Later among the works it cites.
Finite-time analysis of distributed td (0) with linear function approximation on multi-agent reinforcement learning
Doan, T · 2019
Later among the works it cites.
Solving a class of non-convex min-max games using iterative first order methods
Nouiehed, M · 2019
Later among the works it cites.
Efficient algorithms for smooth minimax optimization
Thekumparampil, K. K · 2019
Later among the works it cites.
Variance reduced policy evaluation with smooth function approximation
Wai, H.-T · 2019
Later among the works it cites.
Sharp analysis of simple restarted stochastic gradient methods for min-max optimization
Yan, Y · 2019
Later among the works it cites.