Fetching the paper…
Reading the bibliography…
Off-policy evaluation learns a target policy's value with a historical dataset generated by a different behavior policy.
A class of statistics with asymptotically normal distribution
Hoeffding, W · 1948
Earlier work this paper cites.
On u-statistics and v. mise’statistics for weakly dependent processes
Denker, M. and Keller, G · 1983
Earlier work this paper cites.
Efficient and adaptive estimation for semiparametric models , volume 4
Bickel, P. J., Klaassen, C. A., Bickel, P. J., Ritov, Y., Klaassen, J., Wellner, J. A., and Ritov, Y · 1993
Earlier work this paper cites.
Weak convergence
Van Der Vaart, A. W. and Wellner, J. A · 1996
Earlier work this paper cites.
Empirical likelihood methods with weakly dependent processes
Kitamura, Y. et al · 1997
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D · 2000
Earlier work this paper cites.
Asymptotic statistics , volume 3
Van der Vaart, A. W · 2000
Earlier work this paper cites.
Random forests
Breiman, L · 2001
Earlier work this paper cites.
Marginal mean models for dynamic regimes
Murphy, S. A., van der Laan, M. J., Robins, J. M., and Group, C. P. P. R · 2001
Earlier work this paper cites.
Empirical likelihood
Owen, A. B · 2001
Earlier work this paper cites.
Statistical inference of the value function for reinforcement learning in infinite horizon settings
Shi, C., Zhang, S., Lu, W., and Song, R · 2001
Earlier work this paper cites.
Maximal inequalities and empirical central limit theorems
Dedecker, J. and Louhichi, S · 2002
Earlier work this paper cites.
Basic properties of strong mixing conditions. a survey and some open questions
Bradley, R. C · 2005
Earlier work this paper cites.
Semiparametric theory and missing data
Tsiatis, A · 2007
Earlier work this paper cites.
Self-normalized processes: Limit theory and Statistical Applications
Peña, V. H., Lai, T. L., and Shao, Q.-M · 2008
Earlier work this paper cites.
Higher order influence functions and minimax estimation of nonlinear functionals
Robins, J., Li, L., Tchetgen, E., van der Vaart, A., et al · 2008
Earlier work this paper cites.
Anti-concentration and honest, adaptive confidence bands
Chernozhukov, V., Chetverikov, D., Kato, K., et al · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
Online targeted learning
Van Der Laan, M. J. and Lendle, S. D · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Statistics of robust optimization: A generalized empirical likelihood approach
Duchi, J., Glynn, P., and Namkoong, H · 2016
Cited alongside, same era.
Bootstrapping with models: Confidence intervals for off-policy evaluation
Hanna, J. P., Stone, P., and Niekum, S · 2016
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Jiang, N. and Li, L · 2016
Kallus, N. and Uehara, M · 2019
Later among the works it cites.
Batch policy learning under constraints
Le, H. M., Voloshin, C., and Yue, Y · 2019
Later among the works it cites.
U-statistics: Theory and Practice
Lee, A. J · 2019
Later among the works it cites.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Nachum, O., Chow, Y., Dai, B., and Li, L · 2019
Later among the works it cites.
Doubly robust bias reduction in infinite horizon off-policy estimation
Tang, Z., Feng, Y., Li, L., Zhou, D., and Liu, Q · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Thomas, P. and Brunskill, E · 2016
Cited alongside, same era.
Double/debiased/neyman machine learning of treatment effects
Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., and Newey, W · 2017
Cited alongside, same era.
Evaluating the impact of treating the optimal subgroup
Luedtke, A. R. and van der Laan, M. J · 2017
Cited alongside, same era.
Semiparametric efficient empirical higher order influence function estimators
Mukherjee, R., Newey, W. K., and Robins, J. M · 2017
Cited alongside, same era.
Minimax estimation of a functional on a structured high-dimensional model
Robins, J. M., Li, L., Mukherjee, R., Tchetgen, E. T., van der Vaart, A., et al · 2017
Cited alongside, same era.
Deep reinforcement learning framework for autonomous driving
Sallab, A. E., Abdou, M., Perot, E., and Yogamani, S · 2017
Cited alongside, same era.
A finite time analysis of temporal difference learning with linear function approximation
Bhandari, J., Russo, D., and Singal, R · 2018
Cited alongside, same era.
Uehara, M., Huang, J., and Jiang, N · 2019
Later among the works it cites.
Finite-sample analysis for sarsa with linear function approximation
Zou, S., Xu, T., and Liang, Y · 2019
Later among the works it cites.
Learning dexterous in-hand manipulation
Andrychowicz, O. M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al · 2020
Later among the works it cites.
Coindice: Off-policy confidence interval estimation
Dai, B., Nachum, O., Chow, Y., Li, L., Szepesvari, C., and Schuurmans, D · 2020
Later among the works it cites.
A theoretical analysis of deep q-learning
Fan, J., Wang, Z., Xie, Y., and Yang, Z · 2020
Later among the works it cites.
Accountable off-policy evaluation with kernel bellman statistics
Feng, Y., Ren, T., Tang, Z., and Liu, Q · 2020
Later among the works it cites.
Minimax value interval for off-policy evaluation and policy optimization
Jiang, N. and Huang, J · 2020
Later among the works it cites.
Double reinforcement learning for efficient off-policy evaluation in markov decision processes
Kallus, N. and Uehara, M · 2020
Later among the works it cites.
Batch policy learning in average reward markov decision processes
Liao, P., Qi, Z., and Murphy, S · 2020
Later among the works it cites.
Adaptive estimator selection for off-policy evaluation
Su, Y., Srinath, P., and Krishnamurthy, A · 2020
Later among the works it cites.
Zhang, K. W., Janson, L., and Murphy, S. A · 2020
Later among the works it cites.
Confidence intervals for policy evaluation in adaptive experiments
Hadad, V., Hirshberg, D. A., Zhan, R., Wager, S., and Athey, S · 2021
Closest in time.
Testing mediation effects using logic of boolean matrices
Shi, C. and Li, L · 2021
Closest in time.