Fetching the paper…
Reading the bibliography…
Off-policy evaluation (OPE) is a method for estimating the return of a target policy using some pre-collected observational data generated by a potentially different behavior policy.
Identification and estimation of local average treatment effects
Angrist, J. and G. Imbens (1995) · 1995
Earlier work this paper cites.
Identification of causal effects using instrumental variables
Angrist, J. D., G. W. Imbens, and D. B. Rubin (1996) · 1996
Earlier work this paper cites.
Weak convergence and empirical processes: with applications to statistics
Van Der Vaart, A. W., A. van der Vaart, A. W. van der Vaart, and J. Wellner (1996) · 1996
Earlier work this paper cites.
Gendice: Generalized offline estimation of stationary values
Zhang, R., B. Dai, L. Li, and D. Schuurmans (2020) · 2002
Earlier work this paper cites.
Semiparametric instrumental variable estimation of treatment response models
Abadie, A. (2003) · 2003
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, M. G. and R. Parr (2003) · 2003
Earlier work this paper cites.
Optimal dynamic treatment regimes
Murphy, S. A. (2003) · 2003
Earlier work this paper cites.
Semiparametric theory and missing data
Tsiatis, A. A. (2006) · 2006
Earlier work this paper cites.
Off-policy evaluation in infinite-horizon reinforcement learning with latent confounders
Bennett, A., N. Kallus, L. Li, and A. Mousavi (2021) · 2007
Earlier work this paper cites.
Instrumental variables in models with multiple outcomes: The general unordered case
Heckman, J. J., S. Urzua, and E. Vytlacil (2008) · 2008
Earlier work this paper cites.
An introduction to proximal causal learning
Tchetgen, E. J. T., A. Ying, Y. Cui, X. Shi, and W. Miao (2020) · 2009
Earlier work this paper cites.
Deep jump q-evaluation for offline policy evaluation in continuous action space
Cai, H., C. Shi, R. Song, and W. Lu (2020) · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Li, L., W. Chu, J. Langford, and R. E. Schapire (2010) · 2010
Earlier work this paper cites.
Closing the learning-planning loop with predictive state representations
Boots, B., S. M. Siddiqi, and G. J. Gordon (2011) · 2011
Earlier work this paper cites.
Semiparametric proximal causal inference
Cui, Y., H. Pu, X. Shi, W. Miao, and E. T. Tchetgen (2020) · 2011
Earlier work this paper cites.
Beyond late: Estimation of the average treatment effect with an instrumental variable
Aronow, P. M. and A. Carnegie (2013) · 2013
Earlier work this paper cites.
Alternative identification and inference for the effect of treatment on the treated with an instrumental variable
Tchetgen, E. J. and S. Vansteelandt (2013) · 2013
Earlier work this paper cites.
Tensor decompositions for learning latent variable models
Anandkumar, A., R. Ge, D. Hsu, S. M. Kakade, and M. Telgarsky (2014) · 2014
Earlier work this paper cites.
Gaussian approximation of suprema of empirical processes
Chernozhukov, V., D. Chetverikov, and K. Kato (2014) · 2014
Earlier work this paper cites.
Constructing dynamic treatment regimes in infinite-horizon settings
Ertefaie, A. (2014) · 2014
Earlier work this paper cites.
Bandits with unobserved confounders: A causal approach
Bareinboim, E., A. Forney, and J. Pearl (2015) · 2015
Earlier work this paper cites.
Doubly robust estimation of the local average treatment effect curve
Ogburn, E. L., A. Rotnitzky, and J. M. Robins (2015) · 2015
Earlier work this paper cites.
High-confidence off-policy evaluation
Thomas, P., G. Theocharous, and M. Ghavamzadeh (2015) · 2015
Earlier work this paper cites.
Reinforcement learning of pomdps using spectral methods
Azizzadenesheli, K., A. Lazaric, and A. Anandkumar (2016) · 2016
Earlier work this paper cites.
A pac rl algorithm for episodic pomdps
Guo, Z. D., S. Doroudi, and E. Brunskill (2016) · 2016
Earlier work this paper cites.
Doubly robust off-policy value evaluation for reinforcement learning
Jiang, N. and L. Li (2016) · 2016
Earlier work this paper cites.
Data-efficient off-policy policy evaluation for reinforcement learning
Thomas, P. and E. Brunskill (2016) · 2016
Cited alongside, same era.
Markov decision processes with unobserved confounders: A causal approach
Zhang, J. and E. Bareinboim (2016) · 2016
Cited alongside, same era.
Consistent on-line off-policy evaluation
Hallak, A. and S. Mannor (2017) · 2017
Cited alongside, same era.
Bootstrapping with models: Confidence intervals for off-policy evaluation
Hanna, J. P., P. Stone, and S. Niekum (2017) · 2017
Cited alongside, same era.
Contextual bandits with latent confounders: An nmf approach
Sen, R., K. Shanmugam, M. Kocaoglu, A. Dimakis, and S. Shakkottai (2017) · 2017
Cited alongside, same era.
Econometric methods for program evaluation
Abadie, A. and M. D. Cattaneo (2018) · 2018
Cited alongside, same era.
Off-policy evaluation in partially observable environments
Tennenholtz, G., U. Shalit, and S. Mannor (2020) · 2020
Later among the works it cites.
Minimax weight and q-function learning for off-policy evaluation
Uehara, M., J. Huang, and N. Jiang (2020) · 2020
Later among the works it cites.
Bennett, A. and N. Kallus (2021) · 2021
Later among the works it cites.
Estimating and improving dynamic treatment regimes with a time-varying instrumental variable
Chen, S. and B. Zhang (2021) · 2021
Later among the works it cites.
Bootstrapping fitted q-evaluation for off-policy inference
Hao, B., X. Ji, Y. Duan, H. Lu, C. Szepesvari, and M. Wang (2021) · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
More robust doubly robust off-policy evaluation
Farajtabar, M., Y. Chow, and M. Ghavamzadeh (2018) · 2018
Cited alongside, same era.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Liu, Q., L. Li, Z. Tang, and D. Zhou (2018) · 2018
Cited alongside, same era.
Identifying causal effects with proxy variables of an unmeasured confounder
Miao, W., Z. Geng, and E. J. Tchetgen Tchetgen (2018) · 2018
Cited alongside, same era.
Deep reinforcement learning for vision-based robotic grasping: A simulated comparative evaluation of off-policy methods
Quillen, D., E. Jang, O. Nachum, C. Finn, J. Ibarz, and S. Levine (2018) · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and A. G. Barto (2018) · 2018
Cited alongside, same era.
Bounded, efficient and multiply robust estimation of average treatment effects using instrumental variables
Wang, L. and E. Tchetgen (2018) · 2018
Cited alongside, same era.
Hu, Y. and S. Wager (2021, October) · 2021
Later among the works it cites.
Causal inference under unmeasured confounding with negative controls: A minimax learning approach
Kallus, N., X. Mao, and M. Uehara (2021) · 2021
Later among the works it cites.
Rl for latent mdps: Regret guarantees and a lower bound
Kwon, J., Y. Efroni, C. Caramanis, and S. Mannor (2021) · 2021
Later among the works it cites.
Causal reinforcement learning: An instrumental variable approach
Li, J., Y. Luo, and X. Zhang (2021) · 2021
Later among the works it cites.
Instrumental variable value iteration for causal offline reinforcement learning
Liao, L., Z. Fu, Z. Yang, Y. Wang, M. Kolar, and Z. Wang (2021) · 2021
Later among the works it cites.
Off-policy estimation of long-term average outcomes with applications to mobile health
Liao, P., P. Klasnja, and S. Murphy (2021) · 2021
Later among the works it cites.
A spectral approach to off-policy evaluation for pomdps
Nair, Y. and N. Jiang (2021) · 2021
Later among the works it cites.
Optimal individualized decision rules using instrumental variable methods
Qiu, H., M. Carone, E. Sadikova, M. Petukhova, R. C. Kessler, and A. Luedtke (2021) · 2021
Later among the works it cites.
Deeply-debiased off-policy interval estimation
Shi, C., R. Wan, V. Chernozhukov, and R. Song (2021) · 2021
Later among the works it cites.
Provably efficient causal reinforcement learning with confounded observational data
Wang, L., Z. Yang, and Z. Wang (2021) · 2021
Later among the works it cites.
Deep proxy causal learning and its application to confounded bandit policy evaluation
Xu, L., H. Kanagawa, and A. Gretton (2021) · 2021
Later among the works it cites.
On well-posedness and minimax optimal rates of nonparametric q-function estimation in off-policy evaluation
Chen, X. and Z. Qi (2022) · 2022
Closest in time.
Offline reinforcement learning with instrumental variables in confounded markov decision processes
Fu, Z., Z. Qi, Z. Wang, Z. Yang, Y. Xu, and M. R. Kosorok (2022) · 2022
Closest in time.
Efficiently breaking the curse of horizon in off-policy evaluation with double reinforcement learning
Kallus, N. and M. Uehara (2022) · 2022
Closest in time.
Doubly robust off-policy evaluation for ranking policies under the cascade behavior model
Kiyohara, H., Y. Saito, T. Matsuhiro, Y. Narita, N. Shimizu, and Y. Yamamoto (2022) · 2022
Closest in time.
Batch policy learning in average reward markov decision processes
Liao, P., Z. Qi, R. Wan, P. Klasnja, and S. A. Murphy (2022) · 2022
Closest in time.
Miao, R., Z. Qi, and X. Zhang (2022) · 2022
Closest in time.
A minimax learning approach to off-policy evaluation in confounded partially observable markov decision processes
Shi, C., M. Uehara, J. Huang, and N. Jiang (2022) · 2022
Closest in time.
Off-policy confidence interval estimation with confounded markov decision process
Shi, C., J. Zhu, S. Ye, S. Luo, H. Zhu, and R. Song (2022) · 2022
Closest in time.
Future-dependent value-based off-policy evaluation in pomdps
Uehara, M., H. Kiyohara, A. Bennett, V. Chernozhukov, N. Jiang, N. Kallus, C. Shi, and W. Sun (2022) · 2022
Closest in time.
A review of off-policy evaluation in reinforcement learning
Uehara, M., C. Shi, and N. Kallus (2022) · 2022
Closest in time.