Sample efficient policy search for optimal stopping domains
K. Goel, C. Dann, and E. Brunskill · 2017
Later among the works it cites.
Using options and covariance testing for long horizon off-policy policy evaluation
Z. Guo, P. S. Thomas, and E. Brunskill · 2017
Later among the works it cites.
Bootstrapping with models: Confidence intervals for off-policy evaluation
J. P. Hanna, P. Stone, and S. Niekum · 2017
Later among the works it cites.
Faster rates for policy learning
Original
A. Luedtke and A. Chambaz · 2017
Later among the works it cites.
Quasi-oracle estimation of heterogeneous treatment effects
Original
X. Nie and S. Wager · 2017
Later among the works it cites.
A reinforcement learning approach to weaning of mechanical ventilation in intensive care units
N. Prasad, L.-F. Cheng, C. Chivers, M. Draugelis, and B. E. Engelhardt · 2017
Later among the works it cites.
Reliable decision support using counterfactual models
P. Schulam and S. Saria · 2017
Later among the works it cites.
Residual weighted learning for estimating individualized treatment rules
X. Zhou, N. Mayer-Hamblett, U. Khan, and M. R. Kosorok · 2017
Later among the works it cites.
Design-based analysis in difference-in-differences settings with staggered adoption
Original
S. Athey and G. Imbens · 2018
Later among the works it cites.
More robust doubly robust off-policy evaluation
Original
M. Farajtabar, Y. Chow, and M. Ghavamzadeh · 2018
Later among the works it cites.
Balanced policy evaluation and learning
N. Kallus · 2018
Later among the works it cites.
Confounding-robust policy improvement
N. Kallus and A. Zhou · 2018
Later among the works it cites.
Who should be treated? empirical welfare maximization methods for treatment choice
T. Kitagawa and A. Tetenov · 2018
Later among the works it cites.
Q-learning with nearest neighbors
D. Shah and Q. Xie · 2018
Later among the works it cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
Targeted Learning in Data Science
M. J. Van der Laan and S. Rose · 2018
Later among the works it cites.
Estimation and inference of heterogeneous treatment effects using random forests
S. Wager and S. Athey · 2018
Later among the works it cites.
Interpretable dynamic treatment regimes
Y. Zhang, E. B. Laber, M. Davidian, and A. A. Tsiatis · 2018
Later among the works it cites.
Offline multi-action policy learning: Generalization and optimization
Original
Z. Zhou, S. Athey, and S. Wager · 2018
Later among the works it cites.
Generalized random forests
S. Athey, J. Tibshirani, and S. Wager · 2019
Closest in time.
Information-theoretic considerations in batch reinforcement learning
J. Chen and N. Jiang · 2019
Closest in time.
Double reinforcement learning for efficient off-policy evaluation in markov decision processes
Original
N. Kallus and M. Uehara · 2019
Closest in time.
Metalearners for estimating heterogeneous treatment effects using machine learning
S. R. Künzel, J. S. Sekhon, P. J. Bickel, and B. Yu · 2019
Closest in time.
Batch policy learning under constraints
H. Le, C. Voloshin, and Y. Yue · 2019
Closest in time.
Off-policy policy gradient with state distribution correction
Y. Liu, A. Swaminathan, A. Agarwal, and E. Brunskill · 2019
Closest in time.
Estimating dynamic treatment regimes in mobile health using v-learning
D. J. Luckett, E. B. Laber, A. R. Kahkoska, D. M. Maahs, E. Mayer-Davis, and M. R. Kosorok · 2019
Closest in time.
Dynamic Treatment Regimes: Statistical Methods for Precision Medicine
A. A. Tsiatis, M. Davidian, S. T. Holloway, and E. B. Laber · 2019
Closest in time.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
A. Zanette and E. Brunskill · 2019
Closest in time.
From predictive to prescriptive analytics
D. Bertsimas and N. Kallus · 2020
Closest in time.