Fetching the paper…
Reading the bibliography…
In this work, we consider the problem of estimating a behaviour policy for use in Off-Policy Policy Evaluation (OPE) when the true behaviour policy is unknown.
Approximate nearest neighbors: towards removing the curse of dimensionality
Indyk, Piotr and Motwani, Rajeev · 1998
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, Doina · 2000
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, Damien, Geurts, Pierre, and Wehenkel, Louis · 2005
Earlier work this paper cites.
Doubly Robust Off-policy Evaluation for Reinforcement Learning
Jiang, N. and Li, L · 2015
Earlier work this paper cites.
A Markov Decision Process to suggest optimal treatment of severe infections in intensive care
Komorowski, M., Gordon, A., Celi, L. A., and Faisal, A · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Thomas, Philip and Brunskill, Emma · 2016
Cited alongside, same era.
Importance sampling for fair policy selection
Doroudi, Shayan, Thomas, Philip S, and Brunskill, Emma · 2017
Cited alongside, same era.
On calibration of modern neural networks
Guo, Chuan, Pleiss, Geoff, Sun, Yu, and Weinberger, Kilian Q · 2017
Cited alongside, same era.
Fluid administration in severe sepsis and septic shock, patterns and outcomes: an analysis of a large national database
Marik, Paul E, Linde-Zwirble, Walter T, Bittner, Edward A, Sahatjian, Jennifer, and Hansell, Douglas · 2017
Later among the works it cites.
Continuous state-space models for optimal sepsis treatment-a deep reinforcement learning approach
Raghu, Aniruddh, Komorowski, Matthieu, Celi, Leo Anthony, Szolovits, Peter, and Ghassemi, Marzyeh · 2017
Later among the works it cites.
More robust doubly robust off-policy evaluation
Farajtabar, Mehrdad, Chow, Yinlam, and Ghavamzadeh, Mohammad · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…