Fetching the paper…
Reading the bibliography…
We study the problem of off-policy policy evaluation (OPPE) in RL.
Estimation of regression coefficients when some regressors are not always observed
J. M. Robins, A. Rotnitzky, and L. P. Zhao · 1994
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
D. Precup, R. S. Sutton, and S. P. Singh · 2000
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
M. Kearns and S. Singh · 2002
Earlier work this paper cites.
Clinical data based optimal sti strategies for hiv: a reinforcement learning approach
D. Ernst, G.-B. Stan, J. Goncalves, and L. Wehenkel · 2006
Earlier work this paper cites.
On integral probability metrics, ϕ \phi -divergences and binary classification
B. K. Sriperumbudur, K. Fukumizu, A. Gretton, B. Schölkopf, and G. R. Lanckriet · 2009
Earlier work this paper cites.
Learning bounds for importance weighting
C. Cortes, Y. Mansour, and M. Mohri · 2010
Earlier work this paper cites.
Doubly robust policy evaluation and learning
M. Dudík, J. Langford, and L. Li · 2011
Earlier work this paper cites.
On the empirical estimation of integral probability metrics
B. K. Sriperumbudur, K. Fukumizu, A. Gretton, B. Schölkopf, G. R. Lanckriet, et al · 2012
Earlier work this paper cites.
Offline policy evaluation across representations with applications to educational games
T. Mandel, Y.-E. Liu, S. Levine, E. Brunskill, and Z. Popovic · 2014
Earlier work this paper cites.
High-confidence off-policy evaluation
P. S. Thomas, G. Theocharous, and M. Ghavamzadeh · 2015
Cited alongside, same era.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
N. Jiang and L. Li · 2016
Cited alongside, same era.
Learning representations for counterfactual inference
F. Johansson, U. Shalit, and D. Sontag · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
P. Thomas and E. Brunskill · 2016
Cited alongside, same era.
Bayesian inference of individualized treatment effects using multi-task gaussian processes
A. M. Alaa and M. van der Schaar · 2017
Meta-learners for estimating heterogeneous treatment effects using machine learning
S. Künzel, J. Sekhon, P. Bickel, and B. Yu · 2017
Later among the works it cites.
Reliable decision support using counterfactual models
P. Schulam and S. Saria · 2017
Later among the works it cites.
Estimating individual treatment effect: generalization bounds and algorithms
U. Shalit, F. D. Johansson, and D. Sontag · 2017
Later among the works it cites.
Estimation and inference of heterogeneous treatment effects using random forests
S. Wager and S. Athey · 2017
Later among the works it cites.
Learning optimal policies from observational data
O. Atan, W. R. Zame, and M. van der Schaar · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Using options and covariance testing for long horizon off-policy policy evaluation
Z. Guo, P. S. Thomas, and E. Brunskill · 2017
Cited alongside, same era.
Bootstrapping with models: Confidence intervals for off-policy evaluation
J. P. Hanna, P. Stone, and S. Niekum · 2017
Cited alongside, same era.
M. Farajtabar, Y. Chow, and M. Ghavamzadeh · 2018
Closest in time.
Learning weighted representations for generalization across designs
F. D. Johansson, N. Kallus, U. Shalit, and D. Sontag · 2018
Closest in time.
Ganite: Estimation of individualized treatment effects using generative adversarial nets
J. Yoon, J. Jordon, and M. van der Schaar · 2018
Closest in time.