Fetching the paper…
Reading the bibliography…
We consider off-policy evaluation (OPE) in Partially Observable Markov Decision Processes, where the evaluation policy depends only on observable variables but the behavior policy depends on latent states (Tennenholtz et al.
Eligibility traces for off-policy policy evaluation
Doina Precup, Richard S. Sutton, and Satinder P. Singh · 2000
Earlier work this paper cites.
Matrix Algorithms
G. W. Stewart · 2001
Earlier work this paper cites.
Causality: Models, Reasoning and Inference
Judea Pearl · 2009
Earlier work this paper cites.
Closing the learning-planning loop with predictive state representations
Byron Boots, Sajid M Siddiqi, and Geoffrey J Gordon · 2011
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Miroslav Dudík, John Langford, and Lihong Li · 2011
Earlier work this paper cites.
A spectral algorithm for learning hidden markov models
Daniel Hsu, Sham M. Kakade, and Tong Zhang · 2011
Earlier work this paper cites.
Bias Analysis
Sander Greenland and Timothy Lash · 2012
Earlier work this paper cites.
Measurement bias and effect restoration in causal inference
Manabu Kuroki and Judea Pearl · 2014
Cited alongside, same era.
Spectral learning of predictive state representations with insufficient statistics
Alex Kulesza, Nan Jiang, and Satinder Singh · 2015
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Cited alongside, same era.
Offline evaluation of online reinforcement learning algorithms
Travis Mandel, Yun-En Liu, Emma Brunskill, and Zoran Popović · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip Thomas and Emma Brunskill · 2016
Cited alongside, same era.
Identifying causal effects with proxy variables of an unmeasured confounder
Wang Miao, Zhi Geng, and Eric J Tchetgen Tchetgen · 2018
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Markov decision processes with unobserved confounders : A causal approach
J. Zhang and Elias Bareinboim · 2018
Later among the works it cites.
Counterfactual off-policy evaluation with Gumbel-max structural causal models
Michael Oberst and David Sontag · 2019
Later among the works it cites.
Off-policy evaluation in infinite-horizon reinforcement learning with latent confounders, 2020
Andrew Bennett, Nathan Kallus, Lihong Li, and Ali Mousavi · 2020
Later among the works it cites.
Off-policy evaluation in partially observable environments
Guy Tennenholtz, Uri Shalit, and Shie Mannor · 2020
Later among the works it cites.
Off-policy evaluation in partially observable environments
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Guy Tennenholtz, Uri Shalit, and Shie Mannor · 2020
Later among the works it cites.