Fetching the paper…
Reading the bibliography…
Off-policy evaluation (OPE) is the problem of estimating the value of a target policy using historical data collected under a different logging policy.
Convergence of estimates under dimensionality restrictions
Lucien LeCam · 1973
Earlier work this paper cites.
Multidimensional binary search trees used for associative searching
Jon Louis Bentley · 1975
Earlier work this paper cites.
Optimal global rates of convergence for nonparametric regression
Charles J Stone · 1982
Earlier work this paper cites.
Identifying redundant constraints and implicit equalities in systems of linear constraints
Jan Telgen · 1983
Earlier work this paper cites.
Five balltree construction algorithms
Stephen M Omohundro · 1989
Earlier work this paper cites.
Nonparametric bounds on treatment effects
Charles F Manski · 1990
Earlier work this paper cites.
Confidence intervals for partially identified parameters
Guido W Imbens and Charles F Manski · 2004
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction , volume 2
Trevor Hastie, Robert Tibshirani, Jerome H Friedman, and Jerome H Friedman · 2009
Earlier work this paper cites.
More on confidence intervals for partially identified parameters
Jörg Stoye · 2009
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Earlier work this paper cites.
Partial identification in econometrics
Charles F Manski · 2010
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Miroslav Dudík, John Langford, and Lihong Li · 2011
Earlier work this paper cites.
Geometric approximation algorithms
Sariel Har-Peled · 2011
Earlier work this paper cites.
Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms
Lihong Li, Wei Chu, John Langford, and Xuanhui Wang · 2011
Cited alongside, same era.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Cited alongside, same era.
Webscope dataset R6A: ydata-frontpage-todaymodule-clicks-v1, 2011
Yahoo! · 2011
Cited alongside, same era.
Computational geometry: an introduction
Franco P Preparata and Michael I Shamos · 2012
Cited alongside, same era.
Counterfactual reasoning and learning systems: The example of computational advertising
Léon Bottou, Jonas Peters, Joaquin Quiñonero-Candela, Denis X Charles, D Max Chickering, Elon Portugaly, Dipankar Ray, Patrice Simard, and Ed Snelson · 2013
Cited alongside, same era.
High-dimensional statistics: A non-asymptotic viewpoint , volume 48
Martin J Wainwright · 2019
Later among the works it cites.
Local linear forests
Rina Friedberg, Julie Tibshirani, Susan Athey, and Stefan Wager · 2020
Later among the works it cites.
Minimax value interval for off-policy evaluation and policy optimization
Nan Jiang and Jiawei Huang · 2020
Later among the works it cites.
Off-policy bandits with deficient support
Noveen Sachdeva, Yi Su, and Thorsten Joachims · 2020
Later among the works it cites.
Doubly robust off-policy evaluation with shrinkage
Yi Su, Maria Dimakopoulou, Akshay Krishnamurthy, and Miroslav Dudík · 2020
Later among the works it cites.
Safe policy learning through extrapolation: Application to pre-trial risk assessment
Eli Ben-Michael, D James Greiner, Kosuke Imai, and Zhichao Jiang · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yin Tat Lee and Aaron Sidford · 2015
Cited alongside, same era.
Batch learning from logged bandit feedback through counterfactual risk minimization
Adith Swaminathan and Thorsten Joachims · 2015
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip Thomas and Emma Brunskill · 2016
Cited alongside, same era.
UCI machine learning repository, 2017
Dheeru Dua and Casey Graff · 2017
Cited alongside, same era.
Off-policy evaluation for slate recommendation
Adith Swaminathan, Akshay Krishnamurthy, Alekh Agarwal, Miro Dudik, John Langford, Damien Jose, and Imed Zitouni · 2017
Cited alongside, same era.
Optimal and adaptive off-policy evaluation in contextual bandits
Yu-Xiang Wang, Alekh Agarwal, and Miroslav Dudık · 2017
Cited alongside, same era.
Approximate nearest neighbor search in high dimensions
Alexandr Andoni, Piotr Indyk, and Ilya Razenshteyn · 2018
Cited alongside, same era.
Later among the works it cites.
Off-policy estimation of long-term average outcomes with applications to mobile health
Peng Liao, Predrag Klasnja, and Susan Murphy · 2021
Later among the works it cites.
Off-policy evaluation via adaptive weighting with data from contextual bandits
Ruohan Zhan, Vitor Hadad, David A Hirshberg, and Susan Athey · 2021
Later among the works it cites.
Evaluating stochastic seeding strategies in networks
Alex Chin, Dean Eckles, and Johan Ugander · 2022
Later among the works it cites.
Policy learning with new treatments
Samuel Higbee · 2022
Later among the works it cites.
Using limited trial evidence to choose drug dosage when efficacy and toxicity increase with dose
Charles F Manski · 2023
Closest in time.
Wenlong Mou, Peng Ding, Martin J Wainwright, and Peter L Bartlett · 2023
Closest in time.
Counterfactual evaluation of peer-review assignment policies
Martin Saveski, Steven Jecmen, Nihar B Shah, and Johan Ugander · 2023
Closest in time.