Fetching the paper…
Reading the bibliography…
This paper is concerned with constructing a confidence interval for a target policy's value offline based on a pre-collected observational data in infinite horizon settings.
Efficiently breaking the curse of horizon in off-policy evaluation with double reinforcement learning, arXiv
Kallus, N. and Uehara, M. (2019) · 1909
Earlier work this paper cites.
Efficient and adaptive estimation for semiparametric models
Bickel, P. J., Klaassen, C. A., Bickel, P. J., Ritov, Y., Klaassen, J., Wellner, J. A. and Ritov, Y. (1993) · 1993
Earlier work this paper cites.
Mixture density networks
Bishop, C. (1994) · 1994
Earlier work this paper cites.
Weak convergence and empirical processes, Springer
Van Der Vaart, A. and Wellner, J. (1996) · 1996
Earlier work this paper cites.
Asymptotic statistics
Van der Vaart, A. W. (2000) · 2000
Earlier work this paper cites.
Personalized policy learning using longitudinal mobile health data, arXiv:2001.03258
Hu, X., Qian, M., Cheng, B. and Cheung, Y. K. (2020) · 2001
Earlier work this paper cites.
Li, C., Chan, S. H. and Chen, Y.-T. (2020) · 2003
Earlier work this paper cites.
Optimal dynamic treatment regimes, Journal of the Royal Statistical Society: Series B (Statistical Methodology)
Murphy, S. A. (2003) · 2003
Earlier work this paper cites.
Optimal structural nested models for optimal sequential decisions, Proceedings of the second seattle Symposium in Biostatistics
Robins, J. M. (2004) · 2004
Earlier work this paper cites.
Latent-state models for precision medicine, arXiv:2005.13001
Xu, Z., Laber, E., Staicu, A.-M. and Severus, E. (2020) · 2005
Earlier work this paper cites.
Wang, L., Yang, Z. and Wang, Z. (2020) · 2006
Earlier work this paper cites.
Off-policy evaluation in infinite-horizon reinforcement learning with latent confounders, in
Bennett, A., Kallus, N., Li, L. and Mousavi, A. (2021) · 2007
Earlier work this paper cites.
Batch policy learning in average reward markov decision processes, arXiv:2007.11771
Liao, P., Qi, Z. and Murphy, S. (2020) · 2007
Earlier work this paper cites.
Semiparametric theory and missing data
Tsiatis, A. (2007) · 2007
Earlier work this paper cites.
Accountable off-policy evaluation with kernel bellman statistics, arXiv:2008.06668
Feng, Y., Ren, T., Tang, Z. and Liu, Q. (2020) · 2008
Earlier work this paper cites.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning, Advances in neural information processing systems
Rahimi, A. and Recht, B. (2008) · 2008
Earlier work this paper cites.
Causality
Pearl, J. (2009) · 2009
Earlier work this paper cites.
The economics of two-sided markets, Journal of Economic Perspective
Rysman, M. (2009) · 2009
Earlier work this paper cites.
An introduction to proximal causal learning, arXiv:2009.10982
Tchetgen Tchetgen, E. J., Ying, A., Cui, Y., Shi, X. and Miao, W. (2020) · 2009
Earlier work this paper cites.
Robust batch policy learning in markov decision processes, arXiv:2011.04185
Qi, Z. and Liao, P. (2020) · 2011
Earlier work this paper cites.
Fast nonparametric conditional density estimation, arXiv:1206.5278
Holmes, M. P., Gray, A. G. and Isbell, C. L. (2012) · 2012
Earlier work this paper cites.
A robust method for estimating optimal treatment regimes, Biometrics
Zhang, B., Tsiatis, A. A., Laber, E. B. and Davidian, M. (2012) · 2012
Earlier work this paper cites.
Reinforcement learning in robotics: A survey, The International Journal of Robotics Research
Kober, J., Bagnell, J. A. and Peters, J. (2013) · 2013
Earlier work this paper cites.
Robust estimation of optimal dynamic treatment regimes for sequential treatment decisions, Biometrika
Zhang, B., Tsiatis, A. A., Laber, E. B. and Davidian, M. (2013) · 2013
Earlier work this paper cites.
Dynamic treatment regimes, Annual review of statistics and its application
Chakraborty, B. and Murphy, S. A. (2014) · 2014
Cited alongside, same era.
Gaussian approximation of suprema of empirical processes, The Annals of Statistics
Chernozhukov, V., Chetverikov, D., Kato, K. et al. (2014) · 2014
Cited alongside, same era.
Constructing dynamic treatment regimes in infinite-horizon settings, arXiv:1406.0764
Ertefaie, A. (2014) · 2014
Cited alongside, same era.
Offline policy evaluation across representations with applications to educational games., AAMAS
Mandel, T., Liu, Y.-E., Levine, S., Brunskill, E. and Popovic, Z. (2014) · 2014
Cited alongside, same era.
Evaluating marker-guided treatment selection strategies, Biometrics
Matsouaka, R. A., Li, J. and Cai, T. (2014) · 2014
Cited alongside, same era.
A sparse random projection-based test for overall qualitative treatment effects, Journal of the American Statistical Association
Shi, C., Lu, W. and Song, R. (2019) · 2019
Later among the works it cites.
Dynamic Treatment Regimes: Statistical Methods for Precision Medicine
Tsiatis, A. A., Davidian, M., Holloway, S. T. and Laber, E. B. (2019) · 2019
Later among the works it cites.
Coindice: Off-policy confidence interval estimation, Advances in neural information processing systems
Dai, B., Nachum, O., Chow, Y., Li, L., Szepesvari, C. and Schuurmans, D. (2020) · 2020
Later among the works it cites.
A theoretical analysis of deep q-learning, Learning for Dynamics and Control
Fan, J., Wang, Z., Xie, Y. and Yang, Z. (2020) · 2020
Later among the works it cites.
Robust inference on population indirect causal effects: the generalized front door criterion, Journal of the Royal Statistical Society: Series B (Statistical Methodology)
Fulcher, I. R., Shpitser, I., Marealle, S. and Tchetgen Tchetgen, E. J. (2020) · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
High-confidence off-policy evaluation, Twenty-Ninth AAAI Conference on Artificial Intelligence
Thomas, P. S., Theocharous, G. and Ghavamzadeh, M. (2015) · 2015
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning, International Conference on Machine Learning
Jiang, N. and Li, L. (2016) · 2016
Cited alongside, same era.
Deep reinforcement learning for dialogue generation, arXiv:1606.01541
Li, J., Monroe, W., Ritter, A., Galley, M., Gao, J. and Jurafsky, D. (2016) · 2016
Cited alongside, same era.
Statistical inference for the mean outcome under a possibly non-unique optimal treatment strategy, Annals of statistics
Luedtke, A. R. and Van Der Laan, M. J. (2016) · 2016
Cited alongside, same era.
Markov decision processes with unobserved confounders: A causal approach, Technical report
Zhang, J. and Bareinboim, E. (2016) · 2016
Cited alongside, same era.
Large sample analysis of the median heuristic, arXiv:1707.07269
Garreau, D., Jitkrittum, W. and Kanagawa, M. (2017) · 2017
Cited alongside, same era.
Evaluating the impact of treating the optimal subgroup, Statistical methods in medical research
Luedtke, A. R. and van der Laan, M. J. (2017) · 2017
Cited alongside, same era.
Later among the works it cites.
Confounding-robust policy evaluation in infinite-horizon reinforcement learning, in
Kallus, N. and Zhou, A. (2020) · 2020
Later among the works it cites.
Estimating dynamic treatment regimes in mobile health using v-learning, Journal of the American Statistical Association
Luckett, D. J., Laber, E. B., Kahkoska, A. R., Maahs, D. M., Mayer-Davis, E. and Kosorok, M. R. (2020) · 2020
Later among the works it cites.
Off-policy policy evaluation for sequential decisions under unobserved confounding, Advances in Neural Information Processing Systems
Namkoong, H., Keramati, R., Yadlowsky, S. and Brunskill, E. (2020) · 2020
Later among the works it cites.
Nonparametric regression using deep neural networks with relu activation function, Annals of Statistics
Schmidt-Hieber, J. et al. (2020) · 2020
Later among the works it cites.
Does the markov decision process fit the data: testing for the markov property in sequential decision making, International Conference on Machine Learning
Shi, C., Wan, R., Song, R., Lu, W. and Leng, L. (2020) · 2020
Later among the works it cites.
Multiply robust causal inference with double-negative control adjustment for categorical unmeasured confounding, Journal of the Royal Statistical Society: Series B (Statistical Methodology)
Shi, X., Miao, W., Nelson, J. C. and Tchetgen Tchetgen, E. J. (2020) · 2020
Later among the works it cites.
Off-policy evaluation in partially observable environments., AAAI
Tennenholtz, G., Shalit, U. and Mannor, S. (2020) · 2020
Later among the works it cites.
Minimax weight and q-function learning for off-policy evaluation, in
Uehara, M., Huang, J. and Jiang, N. (2020) · 2020
Later among the works it cites.
Resampling-based confidence intervals for model-free robust inference on optimal treatment regimes, Biometrics
Wu, Y. and Wang, L. (2020) · 2020
Later among the works it cites.
Bennett, A. and Kallus, N. (2021) · 2021
Later among the works it cites.
Causal mediation analysis with mediator values below an assay limit, arXiv:2107.14782
Chernofsky, A., Bosch, R. J. and Lok, J. J. (2021) · 2021
Later among the works it cites.
Bootstrapping statistical inference for off-policy evaluation, arXiv:2102.03607
Hao, B., Ji, X., Duan, Y., Lu, H., Szepesvári, C. and Wang, M. (2021) · 2021
Later among the works it cites.
Kallus, N., Mao, X. and Uehara, M. (2021) · 2021
Later among the works it cites.
Off-policy estimation of long-term average outcomes with applications to mobile health, Journal of the American Statistical Association
Liao, P., Klasnja, P. and Murphy, S. (2021) · 2021
Later among the works it cites.
A spectral approach to off-policy evaluation for pomdps, arXiv:2109.10502
Nair, Y. and Jiang, N. (2021) · 2021
Later among the works it cites.
Deeply-debiased off-policy interval estimation, in
Shi, C., Wan, R., Chernozhukov, V. and Song, R. (2021) · 2021
Later among the works it cites.
Uehara, M., Imaizumi, M., Jiang, N., Kallus, N., Sun, W. and Xie, T. (2021) · 2021
Later among the works it cites.
A minimax learning approach to off-policy evaluation in confounded partially observable markov decision processes, International Conference on Machine Learning
Shi, C., Uehara, M., Huang, J. and Jiang, N. (2022) · 2022
Closest in time.
Statistical inference of the value function for reinforcement learning in infinite horizon settings, Journal of the Royal Statistical Society: Series B (Statistical Methodology)
Shi, C., Zhang, S., Lu, W. and Song, R. (2022) · 2022
Closest in time.