Fetching the paper…
Reading the bibliography…
Markov reward processes (MRPs) are used to model stochastic phenomena arising in operations research, control engineering, robotics, and artificial intelligence, as well as communication and transportation networks.
An Introduction to Probability Theory and its Applications: Volume II
W. Feller · 1966
Earlier work this paper cites.
Problem Complexity and Method Efficiency in Optimization
A. S. Nemirovsky and D. B. Yudin · 1983
Earlier work this paper cites.
Matrix Analysis
R. A. Horn and C. R. Johnson · 1985
Earlier work this paper cites.
Random generation of combinatorial structures from a uniform distribution
M. R. Jerrum, L. G. Valiant, and V. V. Vazirani · 1986
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R. S. Sutton · 1988
Earlier work this paper cites.
A rearrangement inequality and the permutahedron
A. Vince · 1990
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Convergence of stochastic iterative dynamic programming algorithms
T. Jaakkola, M. I. Jordan, and S. P. Singh · 1994
Earlier work this paper cites.
Rates of convergence for empirical processes of stationary mixing sequences
B. Yu · 1994
Earlier work this paper cites.
Dynamic programming and stochastic control
D. P. Bertsekas · 1995
Earlier work this paper cites.
Dynamic programming and stochastic control
D.P. Bertsekas · 1995
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
S. J. Bradtke and A. G. Barto · 1996
Earlier work this paper cites.
Neuro-dynamic programming
D. P. Bertsekas and J. N. Tsitsiklis · 1996
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
J. N. Tsitsiklis and B. Van Roy · 1997
Earlier work this paper cites.
Asynchronous stochastic approximations
V. S. Borkar · 1998
Earlier work this paper cites.
Essentials of stochastic processes
R. Durrett · 1999
Earlier work this paper cites.
Finite-sample convergence rates for Q Q -learning and indirect algorithms
M. Kearns and S. Singh · 1999
Earlier work this paper cites.
The ODE method for convergence of stochastic approximation and reinforcement learning
V. S. Borkar and S. P. Meyn · 2000
Earlier work this paper cites.
Asymptotic statistics
A. W. van der Vaart · 2000
Earlier work this paper cites.
Stability and generalization
O. Bousquet and A. Elisseeff · 2002
Earlier work this paper cites.
Technical update: Least-squares temporal difference learning
J. A. Boyan · 2002
Earlier work this paper cites.
Learning rates for Q Q -learning
E. Even-Dar and Y. Mansour · 2003
Earlier work this paper cites.
On the sample complexity of reinforcement learning
S. M. Kakade · 2003
Earlier work this paper cites.
Inequalities for the L1 deviation of the empirical distribution
T. Weissman, E. Ordentlich, G. Seroussi, S. Verdu, and M. J. Weinberger · 2003
Earlier work this paper cites.
An adaptation theory for nonparametric confidence intervals
T. T. Cai and M. G. Low · 2004
Cited alongside, same era.
On the almost sure rate of convergence of linear stochastic approximation algorithms
V. B. Tadic · 2004
Cited alongside, same era.
Markov decision processes: Discrete stochastic dynamic programming
M. L. Puterman · 2005
Cited alongside, same era.
Reinforcement learning of local shape in the game of Go
D. Silver, R. S. Sutton, and M. Müller · 2007
Cited alongside, same era.
Estimating taxi-out times with a reinforcement learning algorithm
P. Balakrishna, R. Ganesan, L. Sherry, and B. S. Levy · 2008
Cited alongside, same era.
Reinforcement learning in the presence of rare events
J. Frank, S. Mannor, and D. Precup · 2008
Cited alongside, same era.
A finite time analysis of temporal difference learning with linear function approximation
J. Bhandari, D. Russo, and R. Singal · 2018
Later among the works it cites.
Finite sample analyses for TD(0) with function approximation
G. Dalal, B. Szörényi, G. Thoppe, and S. Mannor · 2018
Later among the works it cites.
Open problem: The dependence of sample complexity lower bounds on planning horizon
N. Jiang and A. Agarwal · 2018
Later among the works it cites.
Linear stochastic approximation: How far does constant step-size and iterate averaging go?
C. Lakshminarayanan and C. Szepesvári · 2018
Later among the works it cites.
Implicit regularization in nonconvex statistical estimation: Gradient descent converges linearly for phase retrieval and matrix completion
C. Ma, K. Wang, Y. Chi, and Y. Chen · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A semiparametric statistical approach to model-free policy evaluation
T. Ueno, M. Kawanabe, T. Mori, S. Maeda, and S. Ishii · 2008
Cited alongside, same era.
Reinforcement learning: A tutorial survey and recent advances
A. Gosavi · 2009
Cited alongside, same era.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Cited alongside, same era.
Algorithms for reinforcement learning
C. Szepesvári · 2009
Cited alongside, same era.
Speedy Q Q -learning
M. G. Azar, R. Munos, M. Ghavamzadeh, and H. J. Kappen · 2011
Cited alongside, same era.
Non-asymptotic analysis of stochastic optimization algorithms for machine learning
F. Bach and E. Moulines · 2011
Cited alongside, same era.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
Learning without mixing: Towards a sharp analysis of linear system identification
M. Simchowitz, H. Mania, S. Tu, M. I. Jordan, and B. Recht · 2018
Later among the works it cites.
Near-optimal time and sample complexities for solving Markov decision processes with a generative model
A. Sidford, M. Wang, X. Wu, L. Yang, and Y. Ye · 2018
Later among the works it cites.
Variance reduced value iteration and faster algorithms for solving Markov decision processes
A. Sidford, M. Wang, X. Wu, and Y. Ye · 2018
Later among the works it cites.
Reinforcement learning: Theory and algorithms
A. Agarwal, N. Jiang, and S. M. Kakade · 2019
Closest in time.
Spectral method and regularized mle are both optimal for top- k k ranking
Y. Chen, J. Fan, C. Ma, and K. Wang · 2019
Closest in time.
T. T. Doan, S. T. Maguluri, and J. Romberg · 2019
Closest in time.
Non-asymptotic gap-dependent regret bounds for tabular MDPs
M. Simchowitz and K. G. Jamieson · 2019
Closest in time.
Finite-time error bounds for linear stochastic approximation and TD learning
R. Srikant and L. Ying · 2019
Closest in time.
High-dimensional statistics: A non-asymptotic viewpoint
M. J. Wainwright · 2019
Closest in time.
M. J. Wainwright · 2019
Closest in time.
Variance-reduced Q {Q} -learning is minimax optimal
M. J. Wainwright · 2019
Closest in time.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
A. Zanette and E. Brunskill · 2019
Closest in time.
Almost horizon-free structure-aware best policy identification with a generative model
A. Zanette, M. J. Kochenderfer, and E. Brunskill · 2019
Closest in time.
Model-based reinforcement learning with a generative model is minimax optimal
A. Agarwal, S. M. Kakade, and L. F. Yang · 2020
Closest in time.
Finite-sample analysis of stochastic approximation using smooth convex envelopes
Z. Chen, S. T. Maguluri, S. Shakkottai, and K. Shanmugam · 2020
Closest in time.
Is temporal difference learning optimal? An instance-dependent analysis
K. Khamaru, A. Pananjady, F. Ruan, M. J. Wainwright, and M. I. Jordan · 2020
Closest in time.
Robust machine learning by median-of-means: theory and practice
G. Lecué and M. Lerasle · 2020
Closest in time.
The local geometry of testing in ellipses: Tight control via localized Kolmogorov widths
Y. Wei and M. J. Wainwright · 2020
Closest in time.