Fetching the paper…
Reading the bibliography…
Off-policy reinforcement learning enables near-optimal policy from suboptimal experience, thereby provisions opportunity for artificial intelligence applications in healthcare.
Dynamic Programming
R. Bellman · 1957
Earlier work this paper cites.
The optimal control of partially observable markov processes over the infinite horizon: Discounted costs
E. J. Sondik · 1978
Earlier work this paper cites.
Estimating the dimension of a model
E. S. Gideon · 1978
Earlier work this paper cites.
A drive-reinforcement model of single neuron function: An alternative to the hebbian neuronal model
A. H. Klopf · 1986
Earlier work this paper cites.
Acting optimally in partially observable stochastic domains
A. R. Cassandra, L. P Kaelbling, and M. L. Littman · 1994
Earlier work this paper cites.
Learning policies for partially observable environments: Scaling up
M. L. Littman, A. R. Cassandra, and L. P. Kaelbling · 1995
Earlier work this paper cites.
Bi-pomdp: Bounded, incremental partially-observable markov-model planning
R. Washington · 1997
Earlier work this paper cites.
Reinforcement Learning : An Introduction
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour · 1999
Earlier work this paper cites.
Approximate planning for factored pomdps using belief state simplification
D. A. McAllester and S. Singh · 1999
Earlier work this paper cites.
Value-function approximations for partially observable markov decision processes
M. Hauskrecht · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
D. Precup, R. S. Sutton, and S. P. Singh · 2000
Earlier work this paper cites.
A sparse sampling algorithm for near-optimal planning in large markov decision processes
M. Kearns, Y. Mansour, and A. Y. Ng · 2002
Earlier work this paper cites.
Point-based value iteration: An anytime algorithm for pomdps
J. Pineau, G. J. Gordon, and S. Thrun · 2003
Cited alongside, same era.
A point-based pomdp algorithm for robot planning
M. T. J. Spaan · 2004
Cited alongside, same era.
Heuristic search value iteration for pomdps
T. Smith and R. Simmons · 2004
Cited alongside, same era.
Clinical data based optimal sti strategies for hiv: a reinforcement learning approach
D. Ernst, G. B. Stan, J. Goncalves, and L. Wehenkel · 2006
Cited alongside, same era.
Hybrid pomdp algorithms
S. Paquet, B. Chaib-draa, and S. Ross · 2006
Cited alongside, same era.
Aems: An anytime online search algorithm for approximate policy refinement in large pomdps
S. Ross and B. Chaib-Draa · 2007
Cited alongside, same era.
An application of inverse reinforcement learning to medical records of diabetes treatment
H. Asoh, M. Shiro, S. Akaho, T. Kamishima, K. Hasida, E. Aramaki, and T. Kohro · 2013
Later among the works it cites.
Weighted importance sampling for off-policy learning with linear function approximation
A. R. Mahmood, H. P. van Hasselt, and R. S. Sutton · 2014
Later among the works it cites.
Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs
V. Gulshan, L. Peng, M. Coram, M. C. Stumpe, D. Wu, A. Narayanaswamy, S. Venugopalan, K. Widner, T. Madams, and J. Cuadros · 2016
Later among the works it cites.
Multi-objective markov decision processes for data-driven decision support
D. J. Lizotte and E. B. Laber · 2016
Later among the works it cites.
Optimal medication dosing from suboptimal clinical examples: A deep reinforcement learning approach
S. Nemati, M. M. Ghassemi, and G. D. Clifford · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Informing sequential clinical decision-making through Âreinforcement learning: an empirical study
S. M. Shortreed, E. Laber, D. J. Lizotte, T. S. Stroup, J. Pineau, and S. A. Murphy · 2011
Cited alongside, same era.
Machine learning: a probabilistic perspective
K. P. Murphy · 2012
Cited alongside, same era.
A bayesian view of the poisson-dirichlet process
W. Buntine and M. Hutter · 2012
Cited alongside, same era.
Off-policy actor-critic
T. Degris, M. White, and R. S. Sutton · 2012
Cited alongside, same era.
The use of reinforcement learning algorithms to meet the challenges of an artificial pancreas
M. K. Bothe, L. Dickens, K. Reichel, A. Tellmann, B. Ellger, M. Westphal, and A. A. Faisal · 2013
Cited alongside, same era.
Towards efficient, personalized anesthesia using continuous reinforcement learning for propofol infusion control
C. Lowery and A. A. Faisal · 2013
Cited alongside, same era.
The third international consensus definitions for sepsis and septic shock (sepsis-3)
M. Singer, C.S. Deutschman, C. Seymour, et al · 2016
Later among the works it cites.
Doubly robust off-policy value evaluation for reinforcement learning
N. Jiang and L. Li · 2016
Later among the works it cites.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip S. Thomas and Emma Brunskill · 2016
Later among the works it cites.
Dermatologist-level classification of skin cancer with deep neural networks
A. Esteva, B. Kuprel, R. A. Novoa, J. Ko, S. M. Swetter, H. M. Blau, and S. Thrun · 2017
Later among the works it cites.
A survey on deep learning in medical image analysis
G. Litjens, T. Kooi, B. E. Bejnordi, A. A. A. Setio, F. Ciompi, M. Ghafoorian, J. van der Laak, B. van Ginneken, and C. I Sánchez · 2017
Later among the works it cites.
A reinforcement learning approach to weaning of mechanical ventilation in intensive care units
N. Prasad, L. Cheng, C. Chivers, M. Draugelis, and B. E Engelhardt · 2017
Later among the works it cites.
Mimic-iii, a freely accessible critical care database
A. E. W. Johnson, T. J. Pollard, L. Shen, L. H. Lehman, M. Feng, M. Ghassemi, B. Moody, P. Szolovits, L. Anthony Celi, and R. G. Mark · 2017
Later among the works it cites.