Fetching the paper…
Reading the bibliography…
Off-policy evaluation of sequential decision policies from observational data is necessary in applications of batch reinforcement learning such as education and healthcare.
Finite markov chains. d van nostad co
Kemeny, J. G. and Snell, J. L · 1960
Earlier work this paper cites.
Convex analysis , volume 28
Rockafellar, R. T · 1970
Earlier work this paper cites.
Stability theory for systems of inequalities. part i: Linear systems
Robinson, S. M · 1975
Earlier work this paper cites.
Markov decision problems and state-action frequencies
Altman, E. and Shwartz, A · 1991
Earlier work this paper cites.
Disjunctive programming: Properties of the convex hull of feasible points
Balas, E · 1998
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P., Littman, M. L., and Cassandra, A. R · 1998
Earlier work this paper cites.
Introduction to reinforcement learning , volume 135
Sutton, R. S., Barto, A. G., et al · 1998
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D., Sutton, R., and Singh, S · 2001
Earlier work this paper cites.
S. boyd, l. vanderberghe, 2004
Optimization, C · 2004
Earlier work this paper cites.
Robust dynamic programming
Iyengar, G · 2005
Earlier work this paper cites.
On the empirical state-action frequencies in markov decision processes under general policies
Mannor, S. and Tsitsiklis, J. N · 2005
Earlier work this paper cites.
Robust control of markov decision processes with uncertain transition matrices
Nilim, A. and El Ghaoui, L · 2005
Earlier work this paper cites.
A tutorial on geometric programming
Boyd, S., Kim, S.-J., Vandenberghe, L., and Hassibi, A · 2007
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P · 2010
Earlier work this paper cites.
Interval estimation of population means under unknown but bounded probabilities of sample selection
Aronow, P. and Lee, D · 2012
Cited alongside, same era.
A distributional approach for causal inference using propensity scores
Tan, Z · 2012
Cited alongside, same era.
Robust markov decision processes
Wiesemann, W., Kuhn, D., and Rustem, B · 2013
Cited alongside, same era.
Raam: The benefits of robustness in approximating aggregated mdps in reinforcement learning
Petrik, M. and Subramanian, D · 2014
Cited alongside, same era.
Markov Decision Processes.: Discrete Stochastic Dynamic Programming
Puterman, M. L · 2014
Cited alongside, same era.
High confidence policy improvement
Thomas, P., Theocharous, G., and Ghavamzadeh, M · 2015
Cited alongside, same era.
Interval estimation of individual-level causal effects under unobserved confounding
Kallus, N., Mao, X., and Zhou, A · 2018
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Liu, Q., Li, L., Tang, Z., and Zhou, D · 2018
Later among the works it cites.
Deconfounding reinforcement learning in observational settings
Lu, C., Schölkopf, B., and Hernández-Lobato, J. M · 2018
Later among the works it cites.
Identifying causal effects with proxy variables of an unmeasured confounder
Miao, W., Geng, Z., and Tchetgen Tchetgen, E. J · 2018
Later among the works it cites.
Bounds on the conditional and average treatment effect in the presence of unobserved confounders
Yadlowsky, S., Namkoong, H., Basu, S., Duchi, J., and Tian, L · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Doubly robust off-policy value evaluation for reinforcement learning
Jiang, N. and Li, L · 2016
Cited alongside, same era.
Safe policy improvement by minimizing robust baseline regret
Petrik, M., Ghavamzadeh, M., and Chow, Y · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Thomas, P. and Brunskill, E · 2016
Cited alongside, same era.
Consistent on-line off-policy evaluation
Hallak, A. and Mannor, S · 2017
Cited alongside, same era.
Causal effect inference with deep latent-variable models
Louizos, C., Shalit, U., Mooij, J. M., Sontag, D., Zemel, R., and Welling, M · 2017
Cited alongside, same era.
A reinforcement learning approach to weaning of mechanical ventilation in intensive care units
Prasad, N., Cheng, L.-F., Chivers, C., Draugelis, M., and Engelhardt, B. E · 2017
Cited alongside, same era.
Later among the works it cites.
Sensitivity analysis for unmeasured confounding in coarse structural nested mean models
Yang, S. and Lok, J. J · 2018
Later among the works it cites.
Policy evaluation with latent confounders via optimal balance
Bennett, A. and Kallus, N · 2019
Later among the works it cites.
Off-policy deep reinforcement learning by bootstrapping the covariate shift
Gelada, C. and Bellemare, M. G · 2019
Later among the works it cites.
Guidelines for reinforcement learning in healthcare
Gottesman, O., Johansson, F., Komorowski, M., Faisal, A., Sontag, D., Doshi-Velez, F., and Celi, L. A · 2019
Later among the works it cites.
Kallus, N. and Uehara, M · 2019
Later among the works it cites.
Counterfactual off-policy evaluation with gumbel-max structural causal models
Oberst, M. and Sontag, D · 2019
Later among the works it cites.
Off-policy evaluation in partially observable environments
Tennenholtz, G., Mannor, S., and Shalit, U · 2019
Later among the works it cites.
Near-optimal reinforcement learning in dynamic treatment regimes
Zhang, J. and Bareinboim, E · 2019
Later among the works it cites.