Fetching the paper…
Reading the bibliography…
We offer a theoretical characterization of off-policy evaluation (OPE) in reinforcement learning using function approximation for marginal importance weights and $q$-functions when these are estimated using recent minimax methods.
Orthogonal statistical learning
Foster, D. J. and V. Syrgkanis (2019) · 1901
Earlier work this paper cites.
A unifying approach for doubly-robust ℓ 1 \ell_{1} regularized estimation of causal contrasts
Smucler, E., A. Rotnitzky, and J. M. Robins (2019) · 1904
Earlier work this paper cites.
Kallus, N. and M. Uehara (2019) · 1909
Earlier work this paper cites.
Algaedice: Policy gradient from arbitrary experience
Nachum, O., B. Dai, I. Kostrikov, Y. Chow, L. Li, and D. Schuurmans (2019) · 1912
Earlier work this paper cites.
Improving the sample complexity using global data
Mendelson, S. (2002) · 1991
Earlier work this paper cites.
Large sample estimation and hypothesis testing
Newey, W. K. and D. L. Mcfadden (1994) · 1994
Earlier work this paper cites.
Estimation of regression coefficients when some regressors are not always observed
Robins, J. M., A. Rotnitzky, and L. P. Zhao (1994) · 1994
Earlier work this paper cites.
Asymptotic statistics
van der Vaart, A. W. (1998) · 1998
Earlier work this paper cites.
Sensitivity Analysis for Selection Bias and Unmeasured Confounding in Missing Data and Causal Inference Models
Robins, J. M., A. Rotnitzky, and D. O. Scharfstein (1999) · 1999
Earlier work this paper cites.
Matrix analysis and applied linear algebra
Meyer, C. D. (2000) · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D., R. Sutton, and S. Singh (2000) · 2000
Earlier work this paper cites.
Semi-nonparametric iv estimation of shape-invariant engel curves
Blundell, R., X. Chen, and D. Kristensen (2003) · 2003
Earlier work this paper cites.
Optimal dynamic treatment regimes
Murphy, S. A. (2003) · 2003
Earlier work this paper cites.
Local rademacher complexities
Bartlett, P. L., O. Bousquet, and S. Mendelson (2005, 08) · 2005
Earlier work this paper cites.
Hilbert spaces with applications
Debnath, L. (2005) · 2005
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., P. Geurts, and L. Wehenkel (2005) · 2005
Earlier work this paper cites.
Minimax estimation of conditional moment models
Dikkala, N., G. Lewis, L. Mackey, and V. Syrgkanis (2020) · 2006
Earlier work this paper cites.
Semiparametric Theory and Missing Data
Tsiatis, A. A. (2006) · 2006
Earlier work this paper cites.
Chapter 76 large sample sieve estimation of semi-nonparametric models
Chen, X. (2007) · 2007
Earlier work this paper cites.
Off-policy evaluation via the regularized lagrangian
Yang, M., O. Nachum, B. Dai, L. Li, and D. Schuurmans (2020) · 2007
Earlier work this paper cites.
Near optimal provable uniform convergence in off-policy evaluation for reinforcement learning
Yin, M., Y. Bai, and Y.-X. Wang (2020) · 2007
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
Antos, A., C. Szepesvári, and R. Munos (2008) · 2008
Earlier work this paper cites.
Introduction to Empirical Processes and Semiparametric Inference
Kosorok, M. R. (2008) · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Munos, R. and C. Szepesvári (2008) · 2008
Earlier work this paper cites.
What are the statistical limits of offline rl with linear function approximation?
Wang, R., D. P. Foster, and S. M. Kakade (2020) · 2010
Earlier work this paper cites.
Maximum moment restriction for instrumental variable regression
Zhang, R., M. Imaizumi, B. Schölkopf, and K. Muandet (2020) · 2010
Earlier work this paper cites.
A variant of the wang-foster-kakade lower bound for the discounted setting
Amortila, P., N. Jiang, and T. Xie (2020) · 2011
Earlier work this paper cites.
Sparse feature selection makes batch reinforcement learning more sample efficient
Hao, B., Y. Duan, T. Lattimore, C. Szepesvári, and M. Wang (2020) · 2011
Earlier work this paper cites.
Mathematical statistics: asymptotic Minimax theory
Korostelev, A. P. and O. Korosteleva (2011) · 2011
Earlier work this paper cites.
Statistical methods for dynamic treatment regimes
Chakraborty, B. and E. Moodie (2013) · 2013
Cited alongside, same era.
Policy iteration based on stochastic factorization
Barreto, A. M. S., J. Pineau, and D. Precup (2014) · 2014
Cited alongside, same era.
Off-policy evaluation across representations with applications to educational games
Mandel, T., Y. Liu, S. Levine, E. Brunskill, and Z. Popovic (2014) · 2014
Cited alongside, same era.
Sieve wald and qlr inferences on semi/nonparametric conditional moment models
Chen, X. and D. Pouzo (2015) · 2015
Cited alongside, same era.
Causal inference for statistics, social, and biomedical sciences : an introduction
Imbens, G. (2015) · 2015
Cited alongside, same era.
Adaptive Treatment Strategies in Practice: Planning Trials and Analyzing Data for Personalized Medicine
Kosorok, M. R. and E. E. Moodie (2015) · 2015
Entropy learning for dynamic treatment regimes
Jiang, B., R. Song, J. Li, D. Zeng, W. Lu, X. He, S. Xu, J. Wang, M. Qian, B. Cheng, H. Qiu, A. Luedtke, M. van der Laan, S. Wager, Y. Zhang, E. B. Laber, and N. Kallus (2019) · 2019
Later among the works it cites.
Nonparametric causal effects based on incremental propensity score interventions
Kennedy, E. H. (2019) · 2019
Later among the works it cites.
Precision medicine
Kosorok, M. R. and E. B. Laber (2019) · 2019
Later among the works it cites.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Nachum, O., Y. Chow, B. Dai, and L. Li (2019) · 2019
Later among the works it cites.
Efficient counterfactual learning from bandit feedback
Narita, Y., S. Yasui, and K. Yata (2019) · 2019
Later among the works it cites.
High-Dimensional Statistics : A Non-Asymptotic Viewpoint
Wainwright, M. J. (2019) · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
High confidence policy improvement
Thomas, P., G. Theocharous, and M. Ghavamzadeh (2015) · 2015
Cited alongside, same era.
Regularized policy iteration with nonparametric function spaces
Farahm, A., , M. Ghavamzadeh, C. Szepesvári, and S. Mannor (2016) · 2016
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Jiang, N. and L. Li (2016) · 2016
Cited alongside, same era.
Analysis of classification-based policy iteration algorithms
Lazaric, A., M. Ghavamzadeh, and R. Munos (2016) · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Thomas, P. and E. Brunskill (2016) · 2016
Cited alongside, same era.
Finite-sample optimal estimation and inference on average treatment effects under unconfoundedness
Armstrong, T. B. and M. Kolesár (2017) · 2017
Cited alongside, same era.
Later among the works it cites.
Minimax-optimal off-policy evaluation with linear function approximation
Duan, Y., Z. Jia, and M. Wang (2020) · 2020
Later among the works it cites.
A theoretical analysis of deep q-learning
Fan, J., Z. Wang, Y. Xie, and Z. Yang (2020) · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Z. Yang, Z. Wang, and M. I. Jordan (2020) · 2020
Later among the works it cites.
Morel: Model-based offline reinforcement learning
Kidambi, R., A. Rajeswaran, P. Netrapalli, and T. Joachims (2020) · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., A. Kumar, G. Tucker, and J. Fu (2020) · 2020
Later among the works it cites.
Batch policy learning in average reward markov decision processes
Liao, P., Z. Qi, and S. Murphy (2020) · 2020
Later among the works it cites.
Adaptive approximation and generalization of deep neural network with intrinsic dimensionality
Nakada, R. and M. Imaizumi (2020) · 2020
Later among the works it cites.
Off-policy policy evaluation for sequential decisions under unobserved confounding
Namkoong, H., R. Keramati, S. Yadlowsky, and E. Brunskill (2020) · 2020
Later among the works it cites.
Statistical inference of the value function for reinforcement learning in infinite horizon settings
Shi, C., S. Zhang, W. Lu, and R. Song (2020) · 2020
Later among the works it cites.
Harnessing infinite-horizon off-policy evaluation: Double robustness via duality
Tang, Z., Y. Feng, L. Li, D. Zhou, and Q. Liu (2020) · 2020
Later among the works it cites.
Dynamic treatment regimes : statistical methods for precision medicine
Tsiatis, A. A. A. A. (2020) · 2020
Later among the works it cites.
Minimax weight and q-function learning for off-policy evaluation
Uehara, M., J. Huang, and N. Jiang (2020) · 2020
Later among the works it cites.
Q* approximation schemes for batch reinforcement learning: A theoretical comparison
Xie, T. and N. Jiang (2020) · 2020
Later among the works it cites.
Asymptotically efficient off-policy evaluation for tabular reinforcement learning
Yin, M. and Y.-X. Wang (2020) · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
Yu, T., G. Thomas, L. Yu, S. Ermon, J. Y. Zou, S. Levine, C. Finn, and T. Ma (2020) · 2020
Later among the works it cites.
Gendice: Generalized offline estimation of stationary values
Zhang, R., B. Dai, L. Li, and D. Schuurmans (2020) · 2020
Later among the works it cites.
Sequential causal inference in a single world of connected units
Bibaut, A., M. Petersen, N. Vlassis, M. Dimakopoulou, and M. van der Laan (2021) · 2021
Closest in time.
Risk bounds and rademacher complexity in batch reinforcement learning
Duan, Y., C. Jin, and Z. Li (2021) · 2021
Closest in time.
Off-policy estimation of long-term average outcomes with applications to mobile health
Liao, P., P. Klasnja, and S. Murphy (2021) · 2021
Closest in time.
Minimax off-policy evaluation for multi-armed bandits
Ma, C., B. Zhu, J. Jiao, and M. J. Wainwright (2021) · 2021
Closest in time.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Rashidinejad, P., B. Zhu, C. Ma, J. Jiao, and S. Russell (2021) · 2021
Closest in time.
Deeply-debiased off-policy interval estimation
Shi, C., R. Wan, V. Chernozhukov, and R. Song (2021) · 2021
Closest in time.