Fetching the paper…
Reading the bibliography…
We study the problem of off-policy evaluation (OPE) in reinforcement learning (RL), where the goal is to estimate the performance of a policy from the data generated by another policy(ies).
Some results on generalized difference estimation and generalized regression estimation for finite populations
Cassel, C., Särndal, C., and Wretman, J · 1976
Earlier work this paper cites.
Estimation of regression coefficients when some regressors are not always observed
Robins, J., Rotnitzky, A., and Zhao, L · 1994
Earlier work this paper cites.
Semi-parametric efficiency in multivariate regression models with missing data
Robins, J. and Rotnitzky, A · 1995
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. and Barto, A · 1998
Earlier work this paper cites.
Marginal mean models for dynamic regimes
Murphy, S., van der Laan, M., and Robins, J · 2001
Earlier work this paper cites.
Off-policy temporal difference learning with function approximation
Precup, D., Sutton, R., and Dasgupta, S · 2001
Earlier work this paper cites.
Learning from scarce experience
Peshkin, L. and Shelton, C · 2002
Earlier work this paper cites.
Efficient estimation of average treatment effects using the estimated propensity score
Hirano, K., Imbens, G., and Ridder, W · 2003
Earlier work this paper cites.
Doubly robust estimation in missing data and causal inference models
Bang, H. and Robins, J · 2005
Earlier work this paper cites.
Improving efficiency and robustness of the doubly robust estimator for a population mean with incomplete data
Cao, W., Tsiatis, A., and Davidian, M · 2009
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Dudík, M., Langford, J., and Li, L · 2011
Cited alongside, same era.
Unbiased offline evaluation of contextual bandit-based news article recommendation algorithms
Li, L., an d J. Langford, W. Chu, and Wang, X · 2011
Cited alongside, same era.
Counterfactual reasoning and learning systems: The example of computational advertising
Bottou, L., Peters, J., Quiñonero-Candela, J., Charles, D., Chickering, D. Max, Portugaly, E., Ray, D., Simard, P., and Snelson, E · 2013
Cited alongside, same era.
Off-policy Evaluation in Markov Decision Processes
Paduraru, C · 2013
Cited alongside, same era.
Off-policy learning with eligibility traces: A survey
Geist, M. and Scherrer, B · 2014
Cited alongside, same era.
Weighted importance sampling for off-policy learning with linear function approximation
Mahmood, A., van Hasselt, H., and Sutton, R · 2014
High confidence off-policy evaluation with models
Hanna, J., Stone, P., and Niekum, S · 2016
Later among the works it cites.
Doubly robust off-policy value evaluation for reinforcement learning
Jiang, N. and Li, L · 2016
Later among the works it cites.
Offline evaluation of online reinforcement learning algorithms
Mandel, T., Liu, Y., Brunskill, E., and Popovic, Z · 2016
Later among the works it cites.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M · 2016
Later among the works it cites.
Data-efficient off-policy policy evaluation for reinforcement learning
Thomas, P. and Brunskill, E · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Off-policy evaluation across representations with applications to educational games
Mandel, T., Liu, Y., Levine, S., Brunskill, E., and Popovic, Z · 2014
Cited alongside, same era.
Toward minimax off-policy value estimation
Li, L., Munos, R., and Szepesvàri, Cs · 2015
Cited alongside, same era.
Personalized ad recommendation systems for life-time value optimization with guarantees
Theocharous, G., Thomas, P., and Ghavamzadeh, M · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Eligibility traces for off-policy policy evaluation
Precup, D., Sutton, R., and Singh, S
Cited in the paper.
Eligibility traces for off-policy policy evaluation
Precup, D., Sutton, R., and Singh, S
Cited in the paper.
Gruslys, A., Azar, M., Bellemare, M., and Munos, R · 2017
Later among the works it cites.
Causal effect inference with deep latent-variable models
Louizos, C., Shalit, U., Mooij, J., Sontag, D., Zemel, R., and Welling, M · 2017
Later among the works it cites.
Estimating individual treatment effect: Generalization bounds and algorithms
Shalit, U., Johansson, F., and Sontag, D · 2017
Later among the works it cites.
Off-policy evaluation for slate recommendation
Swaminathan, A., Krishnamurthy, A., Agarwal, A., Dudík, M., Langford, J., Jose, D., and Zitouni, I · 2017
Later among the works it cites.