Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
A kernel loss for solving the Bellman equation
Y. Feng, L. Li, and Q. Liu · 2019
Later among the works it cites.
Efficiently breaking the curse of horizon: Double reinforcement learning in infinite-horizon processes
N. Kallus and M. Uehara · 2019
Later among the works it cites.
High-dimensional statistics: A non-asymptotic viewpoint
M. J. Wainwright · 2019
Later among the works it cites.
Early stopping for kernel boosting algorithms: A general analysis with localized complexities
Y. Wei, F. Yang, and M. J. Wainwright · 2019
Later among the works it cites.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
T. Xie, Y. Ma, and Y.-X. Wang · 2019
Later among the works it cites.
A theoretical analysis of deep Q-learning
J. Fan, Z. Wang, Y. Xie, and Z. Yang · 2020
Later among the works it cites.
Accountable off-policy evaluation with kernel Bellman statistics
Y. Feng, T. Ren, Z. Tang, and Q. Liu · 2020
Later among the works it cites.
Double reinforcement learning for efficient off-policy evaluation in markov decision processes
N. Kallus and M. Uehara · 2020
Later among the works it cites.
Policy evaluation in continuous MDPs with efficient kernelized gradient temporal difference
A. Koppel, G. Warnell, E. Stump, P. Stone, and A. Ribeiro · 2020
Later among the works it cites.
On linear stochastic approximation: Fine-grained Polyak-Ruppert and non-asymptotic concentration
W. Mou, C. J. Li, M. J. Wainwright, P. L. Bartlett, and M. I. Jordan · 2020
Later among the works it cites.
Optimal oracle inequalities for solving projected fixed-point equations
Original
W. Mou, A. Pananjady, and M. J. Wainwright · 2020
Later among the works it cites.
Instance-dependent ℓ ∞ \ell_{\infty} -bounds for policy evaluation in tabular reinforcement learning
A. Pananjady and M. J. Wainwright · 2020
Later among the works it cites.
Asymptotically efficient off-policy evaluation for tabular reinforcement learning
Original
M. Yin and Y.-X. Wang · 2020
Later among the works it cites.
Is temporal difference learning optimal? An instance-dependent analysis
K. Khamaru, A. Pananjady, F. Ruan, M. J. Wainwright, and M. I. Jordan · 2021
Closest in time.
An L 2 {L}^{2} analysis of reinforcement learning in high dimensions with kernel and neural network approximation
Original
J. Long, J. Han, and W. E · 2021
Closest in time.
Sample complexity of offline reinforcement learning with deep relu networks
Original
T. Nguyen-Tang, S. Gupta, H. Tran-The, S. Venkatesh, et al · 2021
Closest in time.