Toward minimax off-policy value estimation
Lihong Li, Rémi Munos, and Csaba Szepesvári · 2015
Cited alongside, same era.
Doubly Robust Off-policy Value Evaluation for Reinforcement Learning
Nan Jiang and Lihong Li · 2016
Cited alongside, same era.
Contextual decision processes with low Bellman rank are PAC-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E. Schapire · 2017
Cited alongside, same era.
Primal-dual π \pi learning: Sample complexity and sublinear run time for ergodic markov decision problems
Original
Mengdi Wang · 2017
Cited alongside, same era.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
CS 598: Notes on Rmax exploration
Nan Jiang · 2018
Cited alongside, same era.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Tengyang Xie, Yifei Ma, and Yu-Xiang Wang · 2019
Cited alongside, same era.
Off-policy policy gradient with state distribution correction
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill · 2019
Cited alongside, same era.
Off-policy estimation of long-term average outcomes with applications to mobile health
Original
Peng Liao, Predrag Klasnja, and Susan Murphy · 2019
Cited alongside, same era.
Gendice: Generalized offline estimation of stationary values
Ruiyi Zhang, Bo Dai, Lihong Li, and Dale Schuurmans · 2019
Cited alongside, same era.
A kernel loss for solving the bellman equation
Yihao Feng, Lihong Li, and Qiang Liu · 2019
Cited alongside, same era.