Fetching the paper…
Reading the bibliography…
Evaluating and optimizing policies in the presence of unobserved confounders is a problem of growing interest in offline reinforcement learning.
Off-policy evaluation in partially observable environments
Tennenholtz, G., Mannor, S., and Shalit, U. (2019) · 1909
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P., Littman, M. L., and Cassandra, A. R. (1998) · 1998
Earlier work this paper cites.
Minimax-optimal off-policy evaluation with linear function approximation
Duan, Y. and Wang, M. (2020) · 2002
Earlier work this paper cites.
Confounding-robust policy evaluation in infinite-horizon reinforcement learning
Kallus, N. and Zhou, A. (2020) · 2002
Earlier work this paper cites.
Off-policy policy evaluation for sequential decisions under unobserved confounding
Namkoong, H., Keramati, R., Yadlowsky, S., and Brunskill, E. (2020) · 2003
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J. (2020) · 2005
Earlier work this paper cites.
Provably efficient causal reinforcement learning with confounded observational data
Wang, L., Yang, Z., and Wang, Z. (2020) · 2006
Earlier work this paper cites.
Off-policy evaluation in infinite-horizon reinforcement learning with latent confounders
Bennett, A., Kallus, N., Li, L., and Mousavi, A. (2020) · 2007
Earlier work this paper cites.
Time-modified Confounding
Platt, R. W., Schisterman, E. F., and Cole, S. R. (2009) · 2009
Earlier work this paper cites.
Methods for dealing with time-dependent confounding
Daniel, R. M., Cousens, S. N., De Stavola, B. L., Kenward, M. G., and Sterne, J. A. C. (2013) · 2013
Earlier work this paper cites.
Off-policy actor-critic
Degris, T., White, M., and Sutton, R. S. (2013) · 2013
Earlier work this paper cites.
Gradient descent converges to minimizers
Lee, J. D., Simchowitz, M., Jordan, M. I., and Recht, B. (2016) · 2016
Cited alongside, same era.
Markov decision processes with unobserved confounders : A causal approach
Zhang, J. and Bareinboim, E. (2016) · 2016
Cited alongside, same era.
Markov chains and mixing times
Levin, D. A. and Peres, Y. (2017) · 2017
Cited alongside, same era.
Handling time varying confounding in observational research
Mansournia, M. A., Etminan, M., Danaei, G., Kaufman, J. S., and Collins, G. (2017) · 2017
Cited alongside, same era.
Causal models adjusting for time-varying confounding—a systematic review of the literature
Clare, P. J., Dobbins, T. A., and Mattick, R. P. (2018) · 2018
Cited alongside, same era.
The limit points of (optimistic) gradient descent in min-max optimization
Daskalakis, C. and Panageas, I. (2018) · 2018
Instrumental variable value iteration for causal offline reinforcement learning
Liao, L., Fu, Z., Yang, Z., Wang, Y., Kolar, M., and Wang, Z. (2021) · 2021
Later among the works it cites.
A spectral approach to off-policy evaluation for pomdps
Nair, Y. and Jiang, N. (2021) · 2021
Later among the works it cites.
A minimax learning approach to off-policy evaluation in confounded partially observable markov decision processes
Shi, C., Uehara, M., Huang, J., and Jiang, N. (2021) · 2021
Later among the works it cites.
Multi-model markov decision processes
Steimle, L., Kaufman, D., and Denton, B. (2021) · 2021
Later among the works it cites.
Offline reinforcement learning with instrumental variables in confounded markov decision processes
Fu, Z., Qi, Z., Wang, Z., Yang, Z., Xu, Y., and Kosorok, M. R. (2022) · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Counterfactual off-policy evaluation with gumbel-max structural causal models
Oberst, M. and Sontag, D. (2019) · 2019
Cited alongside, same era.
Statistically efficient off-policy policy gradients
Kallus, N. and Uehara, M. (2020) · 2020
Cited alongside, same era.
Asymptotically efficient off-policy evaluation for tabular reinforcement learning
Yin, M. and Wang, Y.-X. (2020) · 2020
Cited alongside, same era.
Model-free and model-based policy evaluation when causality is uncertain
Bruns-Smith, D. A. (2021) · 2021
Cited alongside, same era.
Learning mixtures of markov chains and mdps
Kausik, C., Tan, K., and Tewari, A. (2022) · 2022
Closest in time.
Off-policy evaluation for episodic partially observable markov decision processes under non-parametric models
Miao, R., Qi, Z., and Zhang, X. (2022) · 2022
Closest in time.
Future-dependent value-based off-policy evaluation in pomdps
Uehara, M., Kiyohara, H., Bennett, A., Chernozhukov, V., Jiang, N., Kallus, N., Shi, C., and Sun, W. (2022) · 2022
Closest in time.
Off-policy fitted q-evaluation with differentiable function approximators: Z-estimation and inference theory
Zhang, R., Zhang, X., Ni, C., and Wang, M. (2022) · 2022
Closest in time.
Robust fitted-q-evaluation and iteration under sequentially exogenous unobserved confounders
Bruns-Smith, D. and Zhou, A. (2023) · 2023
Closest in time.