Fetching the paper…
Reading the bibliography…
SARSA is an on-policy algorithm to learn a Markov decision process policy in reinforcement learning.
Online Q-learning using connectionist systems
G. A. Rummery and M. Niranjan · 1994
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
S. J. Bradtke and A. G. Barto · 1996
Earlier work this paper cites.
Chattering in SARSA ( λ \lambda )-a CMU learning lab internal report
G. J. Gordon · 1996
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
J. N. Tsitsiklis and B. Roy · 1997
Earlier work this paper cites.
On the existence of fixed points for approximate value iteration and temporal-difference learning
D. P. De Farias and B. Van Roy · 2000
Earlier work this paper cites.
Convergence results for single-step on-policy reinforcement-learning algorithms
S. Singh, T. Jaakkola, M. L. Littman, and C. Szepesvári · 2000
Earlier work this paper cites.
Reinforcement learning with function approximation converges to a region
G. J. Gordon · 2001
Earlier work this paper cites.
Technical update: Least-squares temporal difference learning
J. A. Boyan · 2002
Earlier work this paper cites.
On the existence of fixed points for Q-learning and Sarsa in partially observable domains
T. J. Perkins and M. D. Pendrith · 2002
Earlier work this paper cites.
Least-squares policy iteration
M. G. Lagoudakis and R. Parr · 2003
Earlier work this paper cites.
A convergent form of approximate policy iteration
T. J. Perkins and D. Precup · 2003
Earlier work this paper cites.
Sensitivity and convergence of uniformly ergodic markov chains
A. Y. Mitrophanov · 2005
Earlier work this paper cites.
Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
A. Antos, C. Szepesvari, and R. Munos · 2008
Earlier work this paper cites.
An analysis of reinforcement learning with function approximation
F. S. Melo, S. P. Meyn, and M. I. Ribeiro · 2008
Cited alongside, same era.
Finite-time bounds for fitted value iteration
R. Munos and C. Szepesvari · 2008
Cited alongside, same era.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Cited alongside, same era.
Error propagation for approximate policy and value iteration
A.-M. Farahmand, C. Szepesvari, and R. Munos · 2010
Cited alongside, same era.
LSTD with random projections
M. Ghavamzadeh, A. Lazaric, O. Maillard, and R. Munos · 2010
Cited alongside, same era.
Stochastic approximation: a survey
H. Kushner · 2010
Cited alongside, same era.
On the rate of convergence and error bounds for LSTD ( λ \lambda )
M. Tagorti and B. Scherrer · 2015
Later among the works it cites.
An alternative softmax operator for reinforcement learning
K. Asadi and M. L. Littman · 2016
Later among the works it cites.
Analysis of classification-based policy iteration algorithms
A. Lazaric, M. Ghavamzadeh, and R. Munos · 2016
Later among the works it cites.
A finite time analysis of temporal difference learning with linear function approximation
J. Bhandari, D. Russo, and R. Singal · 2018
Later among the works it cites.
Finite sample analysis of two-timescale stochastic approximation with applications to reinforcement learning
G. Dalal, B. Szorenyi, G. Thoppe, and S. Mannor · 2018
Later among the works it cites.
Finite sample analyses for TD(0) with function approximation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Lazaric, M. Ghavamzadeh, and R. Munos · 2010
Cited alongside, same era.
Dynamic Programming and Optimal Control
D. P. Bertsekas · 2012
Cited alongside, same era.
S. Lacoste-Julien, M. Schmidt, and F. Bach · 2012
Cited alongside, same era.
Finite-sample analysis of least-squares policy iteration
A. Lazaric, M. Ghavamzadeh, and R. Munos · 2012
Cited alongside, same era.
Statistical linear estimation with penalized estimators: An application to reinforcement learning
B. A. Pires and C. Szepesvari · 2012
Cited alongside, same era.
Fast LSTD using stochastic approximation: Finite time analysis and application to traffic control
L. Prashanth, N. Korda, and R. Munos · 2013
Cited alongside, same era.
G. Dalal, B. Szrnyi, G. Thoppe, and S. Mannor · 2018
Later among the works it cites.
Linear stochastic approximation: How far does constant step-size and iterate averaging go?
C. Lakshminarayanan and C. Szepesvari · 2018
Later among the works it cites.
Q-learning with nearest neighbors
D. Shah and Q. Xie · 2018
Later among the works it cites.
Least-squares temporal difference learning for the linear quadratic regulator
S. Tu and B. Recht · 2018
Later among the works it cites.
Finite-time performance bounds and adaptive learning rate selection for two time-scale reinforcement learning
H. Gupta, R. Srikant, and L. Ying · 2019
Closest in time.
Finite-time error bounds for linear stochastic approximation and TD learning
R. Srikant and L. Ying · 2019
Closest in time.
Two time-scale off-policy TD learning: Non-asymptotic analysis over Markovian samples
T. Xu, S. Zou, and Y. Liang · 2019
Closest in time.
A theoretical analysis of deep Q-learning
Z. Yang, Y. Xie, and Z. Wang · 2019
Closest in time.