Fetching the paper…
Reading the bibliography…
We consider evaluating and training a new policy for the evaluation data by using the historical data obtained from a different policy.
A generalization of sampling without replacement from a finite universe
Horvitz, D. G. and Thompson, D. J · 1952
Earlier work this paper cites.
On estimating regression
Nadaraya, E. A · 1964
Earlier work this paper cites.
Smooth regression analysis
Watson, G. S · 1964
Earlier work this paper cites.
Consistent estimation of the influence function of locally asymptotically linear estimators
Klaassen, C. A. J · 1987
Earlier work this paper cites.
Multiple Imputation for Nonresponse in Surveys
Rubin, D. B · 1987
Earlier work this paper cites.
Estimating exposure effects by modelling the expectation of exposure conditional on confounders
Robins, J. M., Mark, S. D., and Newey, W. K · 1992
Earlier work this paper cites.
Nonparametric estimation of mean functionals with data missing at random
Cheng, P. E · 1994
Earlier work this paper cites.
Large sample estimation and hypothesis testing
Newey, W. K. and Mcfadden, D. L · 1994
Earlier work this paper cites.
Estimation of regression coefficients when some regressors are not always observed
Robins, J. M., Rotnitzky, A., and Zhao, L. P · 1994
Earlier work this paper cites.
Toward a curse of dimensionality appropriate (coda) asymptotic theory for semi-parametric models
Robins, J. M. and Ritov, Y. A · 1997
Earlier work this paper cites.
Efficient and Adaptive Estimation for Semiparametric Models
Bickel, P. J., Klaassen, C. A. J., Ritov, Y., and Wellner, J. A · 1998
Earlier work this paper cites.
On the role of the propensity score in efficient semiparametric estimation of average treatment effects
Hahn, J · 1998
Earlier work this paper cites.
Inferences for case-control and semiparametric two-sample density ratio models
Qin, J · 1998
Earlier work this paper cites.
Asymptotic statistics
van der Vaart, A. W · 1998
Earlier work this paper cites.
A matrix extension of the Cauchy-Schwarz inequality
Tripathi, G · 1999
Earlier work this paper cites.
Improving predictive inference under covariate shift by weighting the log-likelihood function
Shimodaira, H · 2000
Earlier work this paper cites.
Asymptotic properties of weighted m -estimators for standard stratified samples
Wooldridge, J. M · 2001
Earlier work this paper cites.
Efficient estimation of average treatment effects using the estimated propensity score
Hirano, K., Imbens, G. W., and Ridder, G · 2003
Earlier work this paper cites.
Semiparametric Theory and Missing Data
Tsiatis, A. A · 2006
Cited alongside, same era.
Direct importance estimation with model selection and its application to covariate shift adaptation
Sugiyama, M., Nakajima, S., Kashima, H., Buenau, P. V., and Kawanabe, M · 2008
Cited alongside, same era.
The offset tree for learning with partial labels
Beygelzimer, A. and Langford, J · 2009
Cited alongside, same era.
Generalizing evidence from randomized clinical trials to target populations
Cole, S. R. and Stuart, E. A · 2010
Cited alongside, same era.
A contextual-bandit approach to personalized news article recommendation
Li, L., Chu, W., Langford, J., and Schapire, R. E · 2010
Cited alongside, same era.
Doubly Robust Policy Evaluation and Learning
Dudík, M., Langford, J., and Li, L · 2011
Cited alongside, same era.
Athey, S. and Wager, S · 2017
Later among the works it cites.
Optimal and adaptive off-policy evaluation in contextual bandits
Wang, Y.-X., Agarwal, A., and Dudik, M · 2017
Later among the works it cites.
Double/debiased machine learning for treatment and structural parameters
Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., and Robins, J · 2018
Later among the works it cites.
More robust doubly robust off-policy evaluation
Farajtabar, M., Chow, Y., and Ghavamzadeh, M · 2018
Later among the works it cites.
Learning weighted representations for generalization across designs
Johansson, F., Kallus, N., Shalit, U., and Sontag, D · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Transportability of causal and statistical relations: A formal approach
Pearl, J. and Bareinboim, E · 2011
Cited alongside, same era.
Cross-validated targeted minimum-loss-based estimation
Zheng, W. and van der Laan, M. J · 2011
Cited alongside, same era.
Statistical analysis of kernel-based least-squares density-ratio estimation
Kanamori, T., Suzuki, T., and Sugiyama, M · 2012
Cited alongside, same era.
Population intervention causal effects based on stochastic interventions
Muñoz, I. D. and Van Der Laan, M · 2012
Cited alongside, same era.
Density Ratio Estimation in Machine Learning
Sugiyama, M., Suzuki, T., and Kanamori, T · 2012
Cited alongside, same era.
Estimating individualized treatment rules using outcome weighted learning
Zhao, Y., Zeng, D., Rush, A. J., and Kosorok, M. R · 2012
Cited alongside, same era.
Who should be treated? empirical welfare maximization methods for treatment choice
Kitagawa, T. and Tetenov, A · 2018
Later among the works it cites.
Offline multi-action policy learning: Generalization and optimization
Zhou, Z., Athey, S., and Wager, S · 2018
Later among the works it cites.
More efficient off-policy evaluation through regularized targeted learning
Bibaut, A., Malenica, I., Vlassis, N., and Van Der Laan, M · 2019
Later among the works it cites.
Semi-parametric efficient policy learning with continuous actions
Chernozhukov, V., Demirer, M., Lewis, G., and Syrgkanis, V · 2019
Later among the works it cites.
Generalizing causal inferences from individuals in randomized trials to all trial-eligible individuals
Dahabreh, I. J., Robertson, S. E., Tchetgen, E. J., Stuart, E. A., and Hernán, M. A · 2019
Later among the works it cites.
Machine learning in the estimation of causal effects: targeted minimum loss-based estimation and double/debiased machine learning
DÃaz, I · 2019
Later among the works it cites.
Intrinsically efficient, stable, and bounded off-policy evaluation for reinforcement learning
Kallus, N. and Uehara, M · 2019
Later among the works it cites.
Nonparametric causal effects based on incremental propensity score interventions
Kennedy, E. H · 2019
Later among the works it cites.
Efficient counterfactual learning from bandit feedback
Narita, Y., Yasui, S., and Yata, K · 2019
Later among the works it cites.
Counterfactual off-policy evaluation with gumbel-max structural causal models
Oberst, M. and Sontag, D · 2019
Later among the works it cites.
High-Dimensional Statistics : A Non-Asymptotic Viewpoint
Wainwright, M. J · 2019
Later among the works it cites.
Double reinforcement learning for efficient off-policy evaluation in markov decision processes
Kallus, N. and Uehara, M · 2020
Closest in time.
Balanced off-policy evaluation in general action spaces
Sondhi, A., Arbour, D., and Dimmery, D · 2020
Closest in time.