Fetching the paper…
Reading the bibliography…
Reinforcement Learning aims at identifying and evaluating efficient control policies from data.
An introduction to the bootstrap
Robert J Tibshirani and Bradley Efron · 1993
Earlier work this paper cites.
Importance sampling for monte carlo estimation of quantiles
Peter W Glynn et al · 1996
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup · 2000
Earlier work this paper cites.
Inductive confidence machines for regression
Harris Papadopoulos, Kostas Proedrou, Volodya Vovk, and Alex Gammerman · 2002
Earlier work this paper cites.
Algorithmic learning in a random world
Vladimir Vovk, Alexander Gammerman, and Glenn Shafer · 2005
Earlier work this paper cites.
Counterfactual reasoning and learning systems: The example of computational advertising
Léon Bottou, Jonas Peters, Joaquin Quiñonero-Candela, Denis X Charles, D Max Chickering, Elon Portugaly, Dipankar Ray, Patrice Simard, and Ed Snelson · 2013
Earlier work this paper cites.
Distribution-free prediction bands for non-parametric regression
Jing Lei and Larry Wasserman · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
High-confidence off-policy evaluation
Philip Thomas, Georgios Theocharous, and Mohammad Ghavamzadeh · 2015
Earlier work this paper cites.
High confidence policy improvement
Philip Thomas, Georgios Theocharous, and Mohammad Ghavamzadeh · 2015
Earlier work this paper cites.
Bootstrapping with models: Confidence intervals for off-policy evaluation
Josiah Hanna, Peter Stone, and Scott Niekum · 2017
Earlier work this paper cites.
Predicting the rate of skin penetration using an aggregated conformal prediction framework
Martin Lindh, Anders Karlén, and Ulf Norinder · 2017
Cited alongside, same era.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou · 2018
Cited alongside, same era.
Batch policy learning under constraints
Hoang Le, Cameron Voloshin, and Yisong Yue · 2019
Cited alongside, same era.
Conformalized quantile regression
Yaniv Romano, Evan Patterson, and Emmanuel Candes · 2019
Cited alongside, same era.
Conformal prediction under covariate shift
Ryan J Tibshirani, Rina Foygel Barber, Emmanuel Candes, and Aaditya Ramdas · 2019
Cited alongside, same era.
Minimax-optimal off-policy evaluation with linear function approximation
Yaqi Duan, Zeyu Jia, and Mengdi Wang · 2020
Cited alongside, same era.
Conformal inference of counterfactuals and individual treatment effects
Lihua Lei and Emmanuel J Candès · 2021
Later among the works it cites.
Deeply-debiased off-policy interval estimation
Chengchun Shi, Runzhe Wan, Victor Chernozhukov, and Rui Song · 2021
Later among the works it cites.
Conformal prediction intervals for markov decision process trajectories
Thomas G Dietterich and Jesse Hostetler · 2022
Later among the works it cites.
Efficiently breaking the curse of horizon in off-policy evaluation with double reinforcement learning
Nathan Kallus and Masatoshi Uehara · 2022
Later among the works it cites.
Safe planning in dynamic environments using conformal prediction
Lars Lindemann, Matthew Cleaveland, Gihyun Shim, and George J Pappas · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Minimax value interval for off-policy evaluation and policy optimization
Nan Jiang and Jiawei Huang · 2020
Cited alongside, same era.
Double reinforcement learning for efficient off-policy evaluation in markov decision processes
Nathan Kallus and Masatoshi Uehara · 2020
Cited alongside, same era.
Application of conformal prediction interval estimations to market makers’ net positions
Wojciech Wisniewski, David Lindsay, and Sian Lindsay · 2020
Cited alongside, same era.
An electronic nose-based assistive diagnostic prototype for lung cancer detection with conformal prediction
Xianghao Zhan, Zhan Wang, Meng Yang, Zhiyuan Luo, You Wang, and Guang Li · 2020
Cited alongside, same era.
Confident off-policy evaluation and selection through self-normalized importance weighting
Ilja Kuzborskij, Claire Vernade, Andras Gyorgy, and Csaba Szepesvári · 2021
Cited alongside, same era.
Charles Lu, Ken Chang, Praveer Singh, and Jayashree Kalpathy-Cramer · 2022
Later among the works it cites.
Awesome conformal prediction, April 2022
Valery Manokhin · 2022
Later among the works it cites.
Statistical inference of the value function for reinforcement learning in infinite-horizon settings
Chengchun Shi, Sheng Zhang, Wenbin Lu, and Rui Song · 2022
Later among the works it cites.
Conformal off-policy prediction in contextual bandits
Muhammad Faaiz Taufiq, Jean-Francois Ton, Rob Cornish, Yee Whye Teh, and Arnaud Doucet · 2022
Later among the works it cites.
A review of off-policy evaluation in reinforcement learning, 2022
Masatoshi Uehara, Chengchun Shi, and Nathan Kallus · 2022
Later among the works it cites.
Conformal prediction for hypersonic flight vehicle classification
Zepu Xi, Xuebin Zhuang, and Hongbo Chen · 2022
Later among the works it cites.