Fetching the paper…
Reading the bibliography…
Many medical decision-making tasks can be framed as partially observed Markov decision processes (POMDPs).
The optimal control of partially observable Markov processes over the infinite horizon: discounted costs
E. J. Sondik · 1978
Earlier work this paper cites.
A tutorial on hidden Markov models and selected applications in speech recognition
L. R. Rabiner · 1989
Earlier work this paper cites.
Reinforcement learning with perceptual aliasing: the perceptual distinctions approach
L. Chrisman · 1992
Earlier work this paper cites.
A note on importance sampling using standardized weights
A. Kong · 1992
Earlier work this paper cites.
An input output HMM architecture
Y. Bengio and P. Frasconi · 1995
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
L. P. Kaelbling, M. L. Littman, and A. R. Cassandra · 1998
Earlier work this paper cites.
Planning treatment of ischemic heart disease with partially observable Markov decision processes
M. Hauskrecht and H. Fraser · 2000
Earlier work this paper cites.
Point-based value iteration: an anytime algorithm for POMDPs
J. Pineau, G. Gordon, and S. Thrun · 2003
Earlier work this paper cites.
Empirical evaluation of the improved rprop learning algorithms
C. Igel and M. Hüsken · 2003
Earlier work this paper cites.
High versus low blood-pressure target in patients with septic shock
P. Asfar, F. Meziani, J. Hamel, F. Grelon, B. Megarbane, N. Anguel, J. Mira, P. Dequin, S. Gergaud, and N. Weiss · 2003
Earlier work this paper cites.
Solving POMDPs with continuous or large discrete observation spaces
J. Hoey and P. Poupart · 2005
Earlier work this paper cites.
Clinical data based optimal STI strategies for HIV: A reinforcement learning approach
D. Ernst, G. Stan, J. Goncalves, and L. Wehenkel · 2006
Earlier work this paper cites.
Using inaccurate models in reinforcement learning
P. Abbeel, M. Quigley, and A. Ng · 2006
Earlier work this paper cites.
Emergency department hypotension predicts sudden unexpected in-hospital mortality: A prospective cohort study
A. E. Jones, V. Yiannibas, C. Johnson, and J. A. Kline · 2006
Earlier work this paper cites.
A reinforcement learning approach for individualizing erythropoietin dosages in hemodialysis patients
J. D. Martín-Guerrero, F. Gomez, E. Soria-Olivas, J. Schmidhuber, M. Climente-Martí, and N. V. Jiménez-Torres · 2009
Cited alongside, same era.
Informing sequential clinical decision-making through reinforcement learning: An empirical study
S. M. Shortreed, E. Laber, D. J. Lizotte, T. S. Stroup, J. Pineau, and S. A. Murphy · 2011
Cited alongside, same era.
Approximate inference for the loss-calibrated Bayesian
S. Lacoste–Julien, F. Huszár, and Z. Ghahramani · 2011
Cited alongside, same era.
Bayesian nonparametric approaches for reinforcement learning in partially observable domains
F. Doshi-Velez · 2012
Cited alongside, same era.
A survey of point-based POMDP solvers
G. Shani, J. Pineau, and R. Kaplow · 2013
Cited alongside, same era.
Guided policy search
S. Levine and V. Koltun · 2013
The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care
M. Komorowski, L. A. Celi, O. Badawi, A. C. Gordon, and A. A. Faisal · 2018
Later among the works it cites.
The actor search tree critic (ASTC) for off-policy POMDP learning in medical decision making
L. Li, M. Komorowski, and A. A. Faisal · 2018
Later among the works it cites.
Improving sepsis treatment strategies by combining deep and kernel-based reinforcement learning
X. Peng, Y. Ding, D. Wihl, O. Gottesman, M. Komorowski, L. H. Lehman, A. Ross, A. Faisal, and F. Doshi-Velez · 2018
Later among the works it cites.
Deep variational reinforcement learning for POMDPs
M. Igl, L. Zintgraf, T. A. Le, F. Wood, and S. Whiteson · 2018
Later among the works it cites.
Iterative value-aware model learning
A. Farahmand · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Effects of fluid administration on arterial load in septic shock patients
M. I. M. García, P. G. González, M. G. Romero, A. G. Cano, C. Oscier, A. Rhodes, R. M. Grounds, and M. Cecconi · 2015
Cited alongside, same era.
Safe reinforcement learning
P. S. Thomas · 2015
Cited alongside, same era.
Robust policy optimization with baseline guarantees
Y. Chow, M. Petrik, and M. Ghavamzadeh · 2015
Cited alongside, same era.
Mimic-iii, a freely accessible critical care database
A. E. W. Johnson, T. J. Pollard, L. Shen, L. H. Lehman, M. Feng, M. Ghassemi, B. Moody, P. Szolovits, L. A. Celi, and R. G. Mark · 2016
Cited alongside, same era.
Continuous state-space models for optimal sepsis treatment: a deep reinforcement learning approach
A. Raghu, M. Komorowski, L. A. Celi, P. Szolovits, and M. Ghassemi · 2017
Cited alongside, same era.
A reinforcement learning approach to weaning of mechanical ventilation in intensive care units
N. Prasad, L. Cheng, C. Chivers, M. Draugelis, and B. E. Engelhardt · 2017
Cited alongside, same era.
A. Raghu, O. Gottesman, Y. Liu, M. Komorowski, A. A. Faisal, F. Doshi-Velez, and E. Brunskill · 2018
Later among the works it cites.
Semi-supervised prediction-constrained topic models
M. C. Hughes, G. Hope, L. Weiner, T. H. Mccoy, R. H. Perlis, E. Sudderth, and F. Doshi-Velez · 2018
Later among the works it cites.
Policy optimization via importance sampling
A. M. Metelli, M. Papini, F. Faccio, and M. Restelli · 2018
Later among the works it cites.
U. M. Girkar, R. Uchimido, L. H. Lehman, P. Szolovits, L. A. Celi, and W. Weng · 2018
Later among the works it cites.
Guidelines for reinforcement learning in healthcare
O. Gottesman, F. Johansson, M. Komorowski, A. Faisal, D. Sontag, F. Doshi-Velez, and L. A. Celi · 2019
Later among the works it cites.
Counterfactual off-policy evaluation with Gumbel-max structural causal models
M. Oberst and D. Sontag · 2019
Later among the works it cites.
Melding the data-decisions pipeline: decision-focused learning for combinatorial optimization
B. Wilder, B. Dilkina, and M. Tambe · 2019
Later among the works it cites.
Off-policy evaluation in partially observable environments
G. Tennenholtz, S. Mannor, and U. Shalit · 2020
Closest in time.