Fetching the paper…
Reading the bibliography…
We consider off-policy evaluation (OPE) in Partially Observable Markov Decision Processes (POMDPs), where the evaluation policy depends only on observable variables and the behavior policy depends on unobservable latent variables.
A kernel loss for solving the bellman equation
Feng, Y., L. Li, and Q. Liu (2019) · 1905
Earlier work this paper cites.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Nachum, O., Y. Chow, B. Dai, and L. Li (2019) · 1906
Earlier work this paper cites.
Kallus, N. and M. Uehara (2019) · 1909
Earlier work this paper cites.
Dual instrumental variable regression
Muandet, K., A. Mehrjou, S. K. Lee, and A. Raj (2019) · 1910
Earlier work this paper cites.
Supervised learning for dynamical system learning
Hefny, A., C. Downey, and G. J. Gordon (2015) · 1971
Earlier work this paper cites.
Asymptotic statistics
van der Vaart, A. W. (1998) · 1998
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D. (2000) · 2000
Earlier work this paper cites.
Popcorn: Partially observed prediction constrained reinforcement learning
Futoma, J., M. C. Hughes, and F. Doshi-Velez (2020) · 2001
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, M. G. and R. Parr (2003) · 2003
Earlier work this paper cites.
Off-policy policy evaluation for sequential decisions under unobserved confounding
Namkoong, H., R. Keramati, S. Yadlowsky, and E. Brunskill (2020) · 2003
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., P. Geurts, and L. Wehenkel (2005) · 2005
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., A. Kumar, G. Tucker, and J. Fu (2020) · 2005
Earlier work this paper cites.
Sample-efficient reinforcement learning of undercomplete pomdps
Jin, C., S. M. Kakade, A. Krishnamurthy, and Q. Liu (2020) · 2006
Earlier work this paper cites.
Provably efficient causal reinforcement learning with confounded observational data
Wang, L., Z. Yang, and Z. Wang (2020) · 2006
Earlier work this paper cites.
Off-policy evaluation in infinite-horizon reinforcement learning with latent confounders
Bennett, A., N. Kallus, L. Li, and A. Mousavi (2021) · 2007
Earlier work this paper cites.
Linear inverse problems in structural econometrics estimation based on spectral decomposition and regularization
Carrasco, M., J.-P. Florens, and E. Renault (2007) · 2007
Earlier work this paper cites.
Batch policy learning in average reward markov decision processes
Liao, P., Z. Qi, and S. Murphy (2020) · 2007
Earlier work this paper cites.
Off-policy evaluation via the regularized lagrangian
Yang, M., O. Nachum, B. Dai, L. Li, and D. Schuurmans (2020) · 2007
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
Antos, A., C. Szepesvári, and R. Munos (2008) · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Munos, R. and C. Szepesvári (2008) · 2008
Earlier work this paper cites.
Coindice: Off-policy confidence interval estimation
Dai, B., O. Nachum, Y. Chow, L. Li, C. Szepesvári, and D. Schuurmans (2020) · 2010
Earlier work this paper cites.
Hilbert space embeddings of hidden markov models
Song, L., B. Boots, S. M. Siddiqi, G. Gordon, and A. Smola (2010) · 2010
Cited alongside, same era.
Closing the learning-planning loop with predictive state representations
Boots, B., S. M. Siddiqi, and G. J. Gordon (2011) · 2011
Cited alongside, same era.
Semiparametric proximal causal inference
Cui, Y., H. Pu, X. Shi, W. Miao, and E. T. Tchetgen (2020) · 2011
Cited alongside, same era.
Cross-validated targeted minimum-loss-based estimation
Zheng, W. and M. J. van der Laan (2011) · 2011
Cited alongside, same era.
The variational method of moments
Bennett, A. and N. Kallus (2020) · 2012
Cited alongside, same era.
A spectral algorithm for learning hidden markov models
Provably efficient reinforcement learning with linear function approximation
Jin, C., Z. Yang, Z. Wang, and M. I. Jordan (2020) · 2020
Later among the works it cites.
Double reinforcement learning for efficient off-policy evaluation in markov decision processes
Kallus, N. and M. Uehara (2020) · 2020
Later among the works it cites.
Confounding-robust policy evaluation in infinite-horizon reinforcement learning
Kallus, N. and A. Zhou (2020) · 2020
Later among the works it cites.
Off-policy evaluation in partially observable environments
Tennenholtz, G., U. Shalit, and S. Mannor (2020) · 2020
Later among the works it cites.
Minimax weight and q-function learning for off-policy evaluation
Uehara, M., J. Huang, and N. Jiang (2020) · 2020
Later among the works it cites.
Asymptotically efficient off-policy evaluation for tabular reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hsu, D., S. M. Kakade, and T. Zhang (2012) · 2012
Cited alongside, same era.
Kernel methods for unobserved confounding: Negative controls, proxies, and instruments
Singh, R. (2020) · 2012
Cited alongside, same era.
Tensor decompositions for learning latent variable models
Anandkumar, A., R. Ge, D. Hsu, S. M. Kakade, and M. Telgarsky (2014) · 2014
Cited alongside, same era.
Spectral learning of predictive state representations with insufficient statistics
Kulesza, A., N. Jiang, and S. Singh (2015) · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al. (2015) · 2015
Cited alongside, same era.
Brockman, G., V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba (2016) · 2016
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Jiang, N. and L. Li (2016) · 2016
Cited alongside, same era.
Yin, M. and Y.-X. Wang (2020) · 2020
Later among the works it cites.
Proximal reinforcement learning: Efficient off-policy evaluation in partially observed markov decision processes
Bennett, A. and N. Kallus (2021) · 2021
Closest in time.
Minimax kernel machine learning for a class of doubly robust functionals
Ghassami, A., A. Ying, I. Shpitser, and E. T. Tchetgen (2021) · 2021
Closest in time.
Confident off-policy evaluation and selection through self-normalized importance weighting
Kuzborskij, I., C. Vernade, A. Gyorgy, and C. Szepesvári (2021) · 2021
Closest in time.
Instrumental variable value iteration for causal offline reinforcement learning
Liao, L., Z. Fu, Z. Yang, Y. Wang, M. Kolar, and Z. Wang (2021) · 2021
Closest in time.
Proximal causal learning with kernels: Two-stage estimation and moment restriction
Mastouri, A., Y. Zhu, L. Gultchin, A. Korba, R. Silva, M. J. Kusner, A. Gretton, and K. Muandet (2021) · 2021
Closest in time.
A spectral approach to off-policy evaluation for pomdps
Nair, Y. and N. Jiang (2021) · 2021
Closest in time.
Instance-dependent l ∞ l_{\infty} -bounds for policy evaluation in tabular reinforcement learning
Pananjady, A. and M. J. Wainwright (2021) · 2021
Closest in time.
Deeply-debiased off-policy interval estimation
Shi, C., R. Wan, V. Chernozhukov, and R. Song (2021) · 2021
Closest in time.
Uehara, M., M. Imaizumi, N. Jiang, N. Kallus, W. Sun, and T. Xie (2021) · 2021
Closest in time.
Deep proxy causal learning and its application to confounded bandit policy evaluation
Xu, L., H. Kanagawa, and A. Gretton (2021) · 2021
Closest in time.
Proximal causal inference for complex longitudinal studies
Ying, A., W. Miao, X. Shi, and E. J. T. Tchetgen (2021) · 2021
Closest in time.
Non-parametric methods for partial identification of causal effects
Zhang, J. and E. Bareinboim (2021) · 2021
Closest in time.
Autoregressive dynamics models for offline policy evaluation and optimization
Zhang, M. R., T. L. Paine, O. Nachum, C. Paduraru, G. Tucker, Z. Wang, and M. Norouzi (2021) · 2021
Closest in time.
Dynamic causal effects evaluation in a/b testing with a reinforcement learning framework
Shi, C., X. Wang, S. Luo, H. Zhu, J. Ye, and R. Song (2022) · 2022
Closest in time.
Off-policy confidence interval estimation with confounded markov decision process
Shi, C., J. Zhu, Y. Shen, S. Luo, H. Zhu, and R. Song (2022) · 2022
Closest in time.