Fetching the paper…
Reading the bibliography…
We propose a new framework for designing estimators for off-policy evaluation in contextual bandits.
Semiparametric efficiency in multivariate regression models with missing data
Robins, J. M. and Rotnitzky, A · 1995
Earlier work this paper cites.
On the role of the propensity score in efficient semiparametric estimation of average treatment effects
Hahn, J · 1998
Earlier work this paper cites.
Efficient estimation of average treatment effects using the estimated propensity score
Hirano, K., Imbens, G. W., and Ridder, G · 2003
Earlier work this paper cites.
Doubly robust estimation in missing data and causal inference models
Bang, H. and Robins, J. M · 2005
Earlier work this paper cites.
Mean-squared-error calculations for average treatment effects
Imbens, G., Newey, W., and Ridder, G · 2007
Earlier work this paper cites.
Demystifying double robustness: A comparison of alternative strategies for estimating a population mean from incomplete data
Kang, J. D., Schafer, J. L., et al · 2007
Earlier work this paper cites.
Online linear optimization and adaptive routing
Awerbuch, B. and Kleinberg, R · 2008
Earlier work this paper cites.
Data-adaptive selection of the truncation level for inverse-probability-of-treatment-weighted estimators
Bembom, O. and van der Laan, M. J · 2008
Earlier work this paper cites.
The price of bandit information for online optimization
Dani, V., Hayes, T. P., and Kakade, S. M · 2008
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
Langford, J. and Zhang, T · 2008
Earlier work this paper cites.
Learning from logged implicit exploration data
Strehl, A., Langford, J., Li, L., and Kakade, S. M · 2010
Cited alongside, same era.
Doubly robust policy evaluation and learning
Dudík, M., Langford, J., and Li, L · 2011
Cited alongside, same era.
Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms
Li, L., Chu, W., Langford, J., and Wang, X · 2011
Cited alongside, same era.
Combinatorial bandits
Cesa-Bianchi, N. and Lugosi, G · 2012
Cited alongside, same era.
Decoupling: from dependence to independence
De la Pena, V. and Giné, E · 2012
Cited alongside, same era.
Counterfactual reasoning and learning systems: The example of computational advertising
Bottou, L., Peters, J., Quiñonero-Candela, J., Charles, D. X., Chickering, D. M., Portugaly, E., Ray, D., Simard, P., and Snelson, E · 2013
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Thomas, P. and Brunskill, E · 2016
Later among the works it cites.
UCI machine learning repository, 2017
Dua, D. and Graff, C · 2017
Later among the works it cites.
A Framework for Optimal Matching for Causal Inference
Kallus, N · 2017
Later among the works it cites.
Off-policy evaluation for slate recommendation
Swaminathan, A., Krishnamurthy, A., Agarwal, A., Dudik, M., Langford, J., Jose, D., and Zitouni, I · 2017
Later among the works it cites.
Optimal and adaptive off-policy evaluation in contextual bandits
Wang, Y.-X., Agarwal, A., and Dudik, M · 2017
Later among the works it cites.
Residual weighted learning for estimating individualized treatment rules
Zhou, X., Mayer-Hamblett, N., Khan, U., and Kosorok, M. R · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Introducing LETOR 4.0 datasets
Qin, T. and Liu, T · 2013
Cited alongside, same era.
Doubly robust policy evaluation and optimization
Dudík, M., Erhan, D., Langford, J., Li, L., et al · 2014
Cited alongside, same era.
Counterfactual estimation and optimization of click metrics in search engines: A case study
Li, L., Chen, S., Kleban, J., and Gupta, A · 2015
Cited alongside, same era.
The value of knowing the propensity score for estimating average treatment effects
Rothe, C · 2016
Cited alongside, same era.
Counterfactual risk minimization: Learning from logged bandit feedback
Swaminathan, A. and Joachims, T
Cited in the paper.
The self-normalized estimator for counterfactual learning
Swaminathan, A. and Joachims, T
Cited in the paper.
Later among the works it cites.
More robust doubly robust off-policy evaluation
Farajtabar, M., Chow, Y., and Ghavamzadeh, M · 2018
Later among the works it cites.
Balanced policy evaluation and learning
Kallus, N · 2018
Later among the works it cites.
Cab: Continuous adaptive blending estimator for policy evaluation and learning
Su, Y., Wang, L., Santacatterina, M., and Joachims, T · 2018
Later among the works it cites.