Policy evaluation and optimization with continuous treatments
Kallus, N. and A. Zhou (2018) · 2018
Later among the works it cites.
Partial mean processes with generated regressors: Continuous treatment effects and nonseparable models
Lee, Y.-Y. (2018) · 2018
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Liu, Q., L. Li, Z. Tang, and D. Zhou (2018) · 2018
Later among the works it cites.
Representation balancing mdps for off-policy policy evaluation
Liu, Y., O. Gottesman, A. Raghu, M. Komorowski, A. A. Faisal, F. Doshi-Velez, and E. Brunskill (2018) · 2018
Later among the works it cites.
Matching on generalized propensity scores with continuous exposures
Original
Wu, X., F. Mealli, M.-A. Kioumourtzoglou, F. Dominici, and D. Braun (2018) · 2018
Later among the works it cites.
More efficient off-policy evaluation through regularized targeted learning
Bibaut, A., I. Malenica, N. Vlassis, and M. Van Der Laan (2019) · 2019
Later among the works it cites.
Double debiased machine learning nonparametric inference with continuous treatments
Colangelo, K. and Y.-Y. Lee (2019) · 2019
Later among the works it cites.
Guidelines for reinforcement learning in healthcare
Gottesman, O., F. Johansson, M. Komorowski, A. Faisal, D. Sontag, F. Doshi-Velez, and L. A. Celi (2019) · 2019
Later among the works it cites.
Causal Inference
Hernan, M. and J. Robins (2019) · 2019
Later among the works it cites.
Batch policy learning under constraints
Le, H., C. Voloshin, and Y. Yue (2019) · 2019
Later among the works it cites.
Non-separable models with high-dimensional data
Su, L., T. Ura, and Y. Zhang (2019) · 2019
Later among the works it cites.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Xie, T., Y. Ma, and Y.-X. Wang (2019) · 2019
Later among the works it cites.
Asymptotically efficient off-policy evaluation for tabular reinforcement learning
Original
Yin, M. and Y.-X. Wang (2020) · 2020
Closest in time.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., D. Meger, and D. Precup (2019) · 2062
Closest in time.