Fetching the paper…
Reading the bibliography…
Off-Policy Evaluation (OPE) serves as one of the cornerstones in Reinforcement Learning (RL).
Optimality and approximation with policy gradient methods in markov decision processes
Alekh Agarwal, Sham M. Kakade, Jason D. Lee, and Gaurav Mahajan · 1908
Earlier work this paper cites.
The bayesian bootstrap
Donald B Rubin · 1981
Earlier work this paper cites.
The jackknife, the bootstrap and other resampling plans
Bradley Efron · 1982
Earlier work this paper cites.
Weak convergence and empirical processes: with applications to statistics
Aad W Van Der Vaart, Aad van der Vaart, Adrianus Willem van der Vaart, and Jon Wellner · 1996
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup · 2000
Earlier work this paper cites.
Asymptotic statistics , volume 3
Aad W Van der Vaart · 2000
Earlier work this paper cites.
Finite time bounds for sampling based fitted value iteration
Csaba Szepesvári and Rémi Munos · 2005
Earlier work this paper cites.
Introduction to empirical processes and semiparametric inference
Michael R Kosorok · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Batch mode reinforcement learning based on the synthesis of artificial trajectories
Raphael Fonteneau, Susan A Murphy, Louis Wehenkel, and Damien Ernst · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Doubly robust off-policy evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2015
Earlier work this paper cites.
Toward minimax off-policy value estimation
Lihong Li, Rémi Munos, and Csaba Szepesvári · 2015
Earlier work this paper cites.
Regularized policy iteration with nonparametric function spaces
Amir-massoud Farahmand, Mohammad Ghavamzadeh, Csaba Szepesvári, and Shie Mannor · 2016
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip Thomas and Emma Brunskill · 2016
Cited alongside, same era.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation, 2019
Statistical bootstrapping for uncertainty estimation in off-policy evaluation
Ilya Kostrikov and Ofir Nachum · 2020
Later among the works it cites.
Minimax weight and q-function learning for off-policy evaluation
Masatoshi Uehara, Jiawei Huang, and Nan Jiang · 2020
Later among the works it cites.
What are the statistical limits of offline rl with linear function approximation?
Ruosong Wang, Dean P Foster, and Sham M Kakade · 2020
Later among the works it cites.
Bridging exploration and general function approximation in reinforcement learning: Provably efficient kernel and neural value iterations
Zhuoran Yang, Chi Jin, Zhaoran Wang, Mengdi Wang, and Michael I Jordan · 2020
Later among the works it cites.
Asymptotically efficient off-policy evaluation for tabular reinforcement learning
Ming Yin and Yu-Xiang Wang · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I. Jordan · 2019
Cited alongside, same era.
Batch policy learning under constraints
Hoang Le, Cameron Voloshin, and Yisong Yue · 2019
Cited alongside, same era.
Optimism in reinforcement learning with generalized linear function approximation
Yining Wang, Ruosong Wang, Simon S Du, and Akshay Krishnamurthy · 2019
Cited alongside, same era.
Tengyang Xie, Yifei Ma, and Yu-Xiang Wang · 2019
Cited alongside, same era.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2020
Cited alongside, same era.
Minimax-optimal off-policy evaluation with linear function approximation
Yaqi Duan, Zeyu Jia, and Mengdi Wang · 2020
Cited alongside, same era.
A theoretical analysis of deep q-learning, 2020
Jianqing Fan, Zhaoran Wang, Yuchen Xie, and Zhuoran Yang · 2020
Cited alongside, same era.
Later among the works it cites.
Optimal policy evaluation using kernel-based temporal difference methods
Yaqi Duan, Mengdi Wang, and Martin J Wainwright · 2021
Later among the works it cites.
Jihao Long, Jiequn Han, and Weinan E · 2021
Later among the works it cites.
Variance-aware off-policy evaluation with linear function approximation
Yifei Min, Tianhao Wang, Dongruo Zhou, and Quanquan Gu · 2021
Later among the works it cites.
Sample complexity of offline reinforcement learning with deep relu networks, 2021
Thanh Nguyen-Tang, Sunil Gupta, Hung Tran-The, and Svetha Venkatesh · 2021
Later among the works it cites.
Statistical inference of the value function for reinforcement learning in infinite horizon settings, 2021
C. Shi, S. Zhang, W. Lu, and R. Song · 2021
Later among the works it cites.
Finite sample analysis of minimax offline reinforcement learning: Completeness, fast rates and first-order efficiency, 2021
Masatoshi Uehara, Masaaki Imaizumi, Nan Jiang, Nathan Kallus, Wen Sun, and Tengyang Xie · 2021
Later among the works it cites.
Provably efficient reinforcement learning for discounted mdps with feature mapping
Dongruo Zhou, Jiafan He, and Quanquan Gu · 2021
Later among the works it cites.
On well-posedness and minimax optimal rates of nonparametric q-function estimation in off-policy evaluation, 2022
Xiaohong Chen and Zhengling Qi · 2022
Closest in time.