Fetching the paper…
Reading the bibliography…
This paper considers batch Reinforcement Learning (RL) with general value function approximation.
Metric entropy of some classes of sets with differentiable boundaries
R. M. Dudley · 1974
Earlier work this paper cites.
Hidden markov models for speech recognition
B. H. Juang and L. R. Rabiner · 1991
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
L. Baird · 1995
Earlier work this paper cites.
Neuro-dynamic programming: an overview
D. P. Bertsekas and J. N. Tsitsiklis · 1995
Earlier work this paper cites.
Support-vector networks
C. Cortes and V. Vapnik · 1995
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
S. J. Bradtke and A. G. Barto · 1996
Earlier work this paper cites.
Experiments with a new boosting algorithm
Y. Freund, R. E. Schapire, et al · 1996
Earlier work this paper cites.
Statistical methods for speech recognition
F. Jelinek · 1997
Earlier work this paper cites.
A brief introduction to boosting
R. E. Schapire · 1999
Earlier work this paper cites.
Least squares support vector machine classifiers
J. A. Suykens and J. Vandewalle · 1999
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
D. Precup · 2000
Earlier work this paper cites.
Bioinformatics: the machine learning approach
P. Baldi, S. Brunak, and F. Bach · 2001
Earlier work this paper cites.
Off-policy temporal-difference learning with function approximation
D. Precup, R. S. Sutton, and S. Dasgupta · 2001
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
P. L. Bartlett and S. Mendelson · 2002
Earlier work this paper cites.
Geometric parameters of kernel machines
S. Mendelson · 2002
Earlier work this paper cites.
Error bounds for approximate policy iteration
R. Munos · 2003
Earlier work this paper cites.
Q* approximation schemes for batch reinforcement learning: A theoretical comparison
T. Xie and N. Jiang · 2003
Earlier work this paper cites.
Local rademacher complexities
P. L. Bartlett, O. Bousquet, S. Mendelson, et al · 2005
Earlier work this paper cites.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
M. Riedmiller · 2005
Earlier work this paper cites.
Convexity, classification, and risk bounds
P. L. Bartlett, M. I. Jordan, and J. D. McAuliffe · 2006
Cited alongside, same era.
Performance bounds in l_p-norm for approximate value iteration
R. Munos · 2007
Cited alongside, same era.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
A. Antos, C. Szepesvári, and R. Munos · 2008
Cited alongside, same era.
Regularized policy iteration
A. M. Farahmand, M. Ghavamzadeh, C. Szepesvári, and S. Mannor · 2008
Cited alongside, same era.
Finite-time bounds for fitted value iteration
R. Munos and C. Szepesvári · 2008
Cited alongside, same era.
Batch value-function approximation with only realizability
T. Xie and N. Jiang · 2008
Probability in high dimension
R. van Handel · 2014
Later among the works it cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Later among the works it cites.
On the uniform convergence of relative frequencies of events to their probabilities
V. N. Vapnik and A. Y. Chervonenkis · 2015
Later among the works it cites.
Regularized policy iteration with nonparametric function spaces
A.-m. Farahmand, M. Ghavamzadeh, C. Szepesvári, and S. Mannor · 2016
Later among the works it cites.
A vector-contraction inequality for rademacher complexities
A. Maurer · 2016
Later among the works it cites.
Contextual decision processes with low bellman rank are pac-learnable
N. Jiang, A. Krishnamurthy, A. Agarwal, J. Langford, and R. E. Schapire · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The elements of statistical learning: data mining, inference, and prediction
T. Hastie, R. Tibshirani, and J. Friedman · 2009
Cited alongside, same era.
Error propagation for approximate policy and value iteration
A. M. Farahmand, R. Munos, and C. Szepesvári · 2010
Cited alongside, same era.
Computer vision: algorithms and applications
R. Szeliski · 2010
Cited alongside, same era.
Introduction to the non-asymptotic analysis of random matrices
R. Vershynin · 2010
Cited alongside, same era.
What are the statistical limits of offline rl with linear function approximation?
R. Wang, D. P. Foster, and S. M. Kakade · 2010
Cited alongside, same era.
Computer vision: a modern approach
D. A. Forsyth and J. Ponce · 2012
Cited alongside, same era.
Later among the works it cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
Information-theoretic considerations in batch reinforcement learning
J. Chen and N. Jiang · 2019
Later among the works it cites.
Batch policy learning under constraints
H. Le, C. Voloshin, and Y. Yue · 2019
Later among the works it cites.
Model-based rl in contextual decision processes: Pac bounds and exponential improvements over model-free approaches
W. Sun, N. Jiang, A. Krishnamurthy, A. Agarwal, and J. Langford · 2019
Later among the works it cites.
High-dimensional statistics: A non-asymptotic viewpoint , volume 48
M. J. Wainwright · 2019
Later among the works it cites.
T. Xie, Y. Ma, and Y.-X. Wang · 2019
Later among the works it cites.
A theoretical analysis of deep q-learning
J. Fan, Z. Wang, Y. Xie, and Z. Yang · 2020
Later among the works it cites.
Double reinforcement learning for efficient off-policy evaluation in markov decision processes
N. Kallus and M. Uehara · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Later among the works it cites.
Minimax weight and q-function learning for off-policy evaluation
M. Uehara, J. Huang, and N. Jiang · 2020
Later among the works it cites.
Near optimal provable uniform convergence in off-policy evaluation for reinforcement learning
M. Yin, Y. Bai, and Y.-X. Wang · 2020
Later among the works it cites.
M. Uehara, M. Imaizumi, N. Jiang, N. Kallus, W. Sun, and T. Xie · 2021
Closest in time.