Fetching the paper…
Reading the bibliography…
Off-policy evaluation (OPE) in both contextual bandits and reinforcement learning allows one to evaluate novel decision policies without needing to conduct exploration, which is often costly or otherwise infeasible.
Large sample estimation and hypothesis testing
W. K. Newey and D. L. Mcfadden · 1994
Earlier work this paper cites.
Estimation of regression coefficients when some regressors are not always observed
J. M. Robins, A. Rotnitzky, and L. P. Zhao · 1994
Earlier work this paper cites.
Asymptotic statistics
A. W. van der Vaart · 1998
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
D. Precup, R. Sutton, and S. Singh · 2000
Earlier work this paper cites.
Empirical likelihood
A. Owen · 2001
Earlier work this paper cites.
Optimal dynamic treatment regimes
S. A. Murphy · 2003
Earlier work this paper cites.
On likelihood approach for monte carlo integration
Z. Tan · 2004
Earlier work this paper cites.
A distributional approach for causal inference using propensity scores
Z. Tan · 2006
Earlier work this paper cites.
Semiparametric Theory and Missing Data
A. Tsiatis · 2006
Earlier work this paper cites.
Demystifying double robustness: A comparison of alternative strategies for estimating a population mean from incomplete data
J. D. Y. Kang and J. L. Schafer · 2007
Earlier work this paper cites.
Comment: Performance of double-robust estimators when "inverse probability" weights are highly variable
J. Robins, M. Sued, Q. Lei-Gomez, and A. Rotnitzky · 2007
Cited alongside, same era.
Empirical efficiency maximization: Improved locally efficient covariate adjustment in randmized experiments and survival analysis
D. B. Rubin and M. J. V. der Laan · 2008
Cited alongside, same era.
Improving efficiency and robustness of the doubly robust estimator for a population mean with incomplete data
W. Cao, A. A. Tsiatis, and M. Davidian · 2009
Cited alongside, same era.
Bounded, efficient and doubly robust estimation with inverse weighting
Z. Tan · 2010
Cited alongside, same era.
Identifying sources of variation and the flow of information in biochemical networks
C. G. Bowsher and P. S. Swain · 2012
Cited alongside, same era.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, , and W. Zaremba · 2016
Later among the works it cites.
Doubly robust off-policy value evaluation for reinforcement learning
N. Jiang and L. Li · 2016
Later among the works it cites.
Safe and efficient off-policy reinforcement learning
R. Munos, T. Stepleton, A. Harutyunyan, and M. Bellemare · 2016
Later among the works it cites.
Data-efficient off-policy policy evaluation for reinforcement learning
P. Thomas and E. Brunskill · 2016
Later among the works it cites.
Optimal and adaptive off-policy evaluation in contextual bandits
Y.-X. Wang, A. Agarwal, and M. Dudik · 2017
Later among the works it cites.
More robust doubly robust off-policy evaluation
M. Farajtabar, Y. Chow, and M. Ghavamzadeh · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Dudík, D. Erhan, J. Langford, and L. Li · 2014
Cited alongside, same era.
Weighted importance sampling for off-policy learning with linear function approximation
A. R. Mahmood, H. P. van Hasselt, and R. S. Sutton · 2014
Cited alongside, same era.
Off-policy evaluation across representations with applications to educational games
T. Mandel, Y. Liu, S. Levine, E. Brunskill, and Z. Popovic · 2014
Cited alongside, same era.
Toward minimax off-policy value estimation
L. Li, R. Munos, and C. Szepesvari · 2015
Cited alongside, same era.
The self-normalized estimator for counterfactual learning
A. Swaminathan and T. Joachims · 2015
Cited alongside, same era.
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Q. Liu, L. Li, Z. Tang, and D. Zhou · 2018
Later among the works it cites.
Reinforcement learning : an introduction
R. S. Sutton · 2018
Later among the works it cites.
Efficient counterfactual learning from bandit feedback
Y. Narita, S. Yasui, and K. Yata · 2019
Closest in time.