Fetching the paper…
Reading the bibliography…
As a subfield of machine learning, reinforcement learning (RL) aims at empowering one's capabilities in behavioural decision making by using interaction experience with the world and an evaluative feedback.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in ICML , 2016, pp. 1928–1937
1937
Earlier work this paper cites.
R. W. Jelliffe, J. Buell, R. Kalaba, R. Sridhar, and R. Rockwell, “A computer program for digitalis dosage regimens,” Mathematical Biosciences , vol. 9, pp. 179–193, 1970
1970
Earlier work this paper cites.
A. M. Albisser, B. Leibel, T. Ewart, Z. Davidovac, C. Botz, W. Zingg, H. Schipper, and R. Gander, “Clinical control of diabetes by the artificial pancreas,” Diabetes , vol. 23, no. 5, pp. 397–404, 1974
1974
Earlier work this paper cites.
R. N. Bergman, Y. Z. Ider, C. R. Bowden, and C. Cobelli, “Quantitative estimation of insulin sensitivity.” American Journal of Physiology-Endocrinology And Metabolism , vol. 236, no. 6, p. E667, 1979
1979
Earlier work this paper cites.
R. E. Bellman, Mathematical methods in medicine . World Scientific Publishing Co., Inc., 1983
1983
Earlier work this paper cites.
J. Pazis and R. Parr, “Efficient pac-optimal exploration in concurrent, continuous state mdps with delayed updates.” in AAAI , 2016, pp. 1977–1985
1985
Earlier work this paper cites.
C. J. Watkins and P. Dayan, “Q-learning,” Machine Learning , vol. 8, no. 3-4, pp. 279–292, 1992
1992
Earlier work this paper cites.
G. A. Rummery and M. Niranjan, On-line Q-learning using connectionist systems . University of Cambridge, Department of Engineering Cambridge, England, 1994, vol. 37
1994
Earlier work this paper cites.
C. Hu, W. S. Lovejoy, and S. L. Shafer, “Comparison of some control strategies for three-compartment pk/pd models,” Journal of Pharmacokinetics and Biopharmaceutics , vol. 22, no. 6, pp. 525–550, 1994
1994
Earlier work this paper cites.
T. Jaakkola, S. P. Singh, and M. I. Jordan, “Reinforcement learning algorithm for partially observable markov decision problems,” in Advances in Neural Information Processing Systems , 1995, pp. 345–352
1995
Earlier work this paper cites.
V. Vapnik, S. E. Golowich, and A. J. Smola, “Support vector method for function approximation, regression estimation and signal processing,” in Advances in Neural Information Processing Systems , 1997, pp. 281–287
1997
Earlier work this paper cites.
R. Ballard-Barbash, S. H. Taplin, B. C. Yankaskas, V. L. Ernster, R. D. Rosenberg, P. A. Carney, W. E. Barlow, B. M. Geller, K. Kerlikowske, B. K. Edwards et al. , “Breast cancer surveillance consortium: a national mammography screening and outcomes database.” American Journal of Roentgenology , vol. 169, no. 4, pp. 1001–1008, 1997
1997
Earlier work this paper cites.
M. Kearns and D. Koller, “Efficient reinforcement learning in factored mdps,” in IJCAI , vol. 16, 1999, pp. 740–747
1999
Earlier work this paper cites.
A. Y. Ng, D. Harada, and S. Russell, “Policy invariance under reward transformations: Theory and application to reward shaping,” in ICML , vol. 99, 1999, pp. 278–287
1999
Earlier work this paper cites.
A. Y. Ng, S. J. Russell et al. , “Algorithms for inverse reinforcement learning.” in ICML , 2000, pp. 663–670
2000
Earlier work this paper cites.
E. H. Wagner, B. T. Austin, C. Davis, M. Hindmarsh, J. Schaefer, and A. Bonomi, “Improving chronic illness care: translating evidence into action,” Health Affairs , vol. 20, no. 6, pp. 64–78, 2001
2001
Earlier work this paper cites.
R. I. Brafman and M. Tennenholtz, “R-max-a general polynomial time algorithm for near-optimal reinforcement learning,” Journal of Machine Learning Research , vol. 3, no. Oct, pp. 213–231, 2002
2002
Earlier work this paper cites.
C. Guestrin, R. Patrascu, and D. Schuurmans, “Algorithm-directed exploration for model-based reinforcement learning in factored mdps,” in ICML , 2002, pp. 235–242
2002
Earlier work this paper cites.
J. K. Lunceford, M. Davidian, and A. A. Tsiatis, “Estimation of survival distributions of treatment policies in two-stage randomization designs in clinical trials,” Biometrics , vol. 58, no. 1, pp. 48–57, 2002
2002
Earlier work this paper cites.
D. Ormoneit and Ś. Sen, “Kernel-based reinforcement learning,” Machine learning , vol. 49, no. 2-3, pp. 161–178, 2002
2002
Earlier work this paper cites.
M. Kearns and S. Singh, “Near-optimal reinforcement learning in polynomial time,” Machine learning , vol. 49, no. 2-3, pp. 209–232, 2002
2002
Earlier work this paper cites.
S. M. Kakade et al. , “On the sample complexity of reinforcement learning,” Ph.D. dissertation, University of London London, England, 2003
2003
Earlier work this paper cites.
M. G. Lagoudakis and R. Parr, “Least-squares policy iteration,” Journal of Machine Learning Research , vol. 4, no. Dec, pp. 1107–1149, 2003
2003
Earlier work this paper cites.
C. Guestrin, D. Koller, R. Parr, and S. Venkataraman, “Efficient solution algorithms for factored mdps,” Journal of Artificial Intelligence Research , vol. 19, pp. 399–468, 2003
2003
Earlier work this paper cites.
A. G. Barto and S. Mahadevan, “Recent advances in hierarchical reinforcement learning,” Discrete Event Dynamic Systems , vol. 13, no. 1-2, pp. 41–77, 2003
2003
Earlier work this paper cites.
L. G. De Pillis and A. Radunskaya, “The dynamics of an optimally controlled tumor model: A case study,” Mathematical and Computer Modelling , vol. 37, no. 11, pp. 1221–1244, 2003
2003
Earlier work this paper cites.
S. A. Murphy, “Optimal dynamic treatment regimes,” Journal of the Royal Statistical Society: Series B (Statistical Methodology) , vol. 65, no. 2, pp. 331–355, 2003
2003
Earlier work this paper cites.
Z. Wang, T. Schaul, M. Hessel, H. Hasselt, M. Lanctot, and N. Freitas, “Dueling network architectures for deep reinforcement learning,” in International Conference on Machine Learning , 2016, pp. 1995–2003
2003
Earlier work this paper cites.
P. Abbeel and A. Y. Ng, “Apprenticeship learning via inverse reinforcement learning,” in Proceedings of the twenty-first international conference on Machine learning . ACM, 2004, p. 1
2004
Earlier work this paper cites.
R. Hovorka, V. Canonico, L. J. Chassin, U. Haueter, M. Massi-Benedetti, M. O. Federici, T. R. Pieber, H. C. Schaller, L. Schaupp, T. Vering et al. , “Nonlinear model predictive control of glucose concentration in subjects with type 1 diabetes,” Physiological Measurement , vol. 25, no. 4, p. 905, 2004
2004
Earlier work this paper cites.
B. M. Adams, H. T. Banks, H.-D. Kwon, and H. T. Tran, “Dynamic multidrug therapies for hiv: Optimal and sti control approaches,” Mathematical Biosciences and Engineering , vol. 1, no. 2, pp. 223–241, 2004
2004
Earlier work this paper cites.
A. J. Rush, M. Fava, S. R. Wisniewski, P. W. Lavori, M. H. Trivedi, H. A. Sackeim, M. E. Thase, A. A. Nierenberg, F. M. Quitkin, T. M. Kashner et al. , “Sequenced treatment alternatives to relieve depression (star* d): rationale and design,” Controlled clinical trials , vol. 25, no. 1, pp. 119–142, 2004
2004
Earlier work this paper cites.
B. L. Moore, E. D. Sinzinger, T. M. Quasny, and L. D. Pyeatt, “Intelligent control of closed-loop sedation in simulated icu patients.” in FLAIRS Conference , 2004, pp. 109–114
2004
Earlier work this paper cites.
G. W. Taylor, “A reinforcement learning framework for parameter control in computer vision applications,” in Computer and Robot Vision, 2004. Proceedings. First Canadian Conference on . IEEE, 2004, pp. 496–503
2004
Earlier work this paper cites.
I. Kola and J. Landis, “Can the pharmaceutical industry reduce attrition rates?” Nature reviews Drug discovery , vol. 3, no. 8, p. 711, 2004
2004
Earlier work this paper cites.
M. Riedmiller, “Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method,” in European Conference on Machine Learning . Springer, 2005, pp. 317–328
2005
Earlier work this paper cites.
D. Ernst, P. Geurts, and L. Wehenkel, “Tree-based batch mode reinforcement learning,” Journal of Machine Learning Research , vol. 6, no. Apr, pp. 503–556, 2005
2005
Earlier work this paper cites.
A. J. Schaefer, M. D. Bailey, S. M. Shechter, and M. S. Roberts, “Modeling medical treatment using markov decision processes,” in Operations Research and Health Care . Springer, 2005, pp. 593–612
2005
Earlier work this paper cites.
S. A. Murphy, “An experimental design for the development of adaptive treatment strategies,” Statistics in Medicine , vol. 24, no. 10, pp. 1455–1481, 2005
2005
Earlier work this paper cites.
W. H. Organization, Preventing chronic diseases: a vital investment . World Health Organization, 2005
2005
Earlier work this paper cites.
B. W. Bequette, “A critical assessment of algorithms and challenges in the development of a closed-loop artificial pancreas,” Diabetes Technology & Therapeutics , vol. 7, no. 1, pp. 28–47, 2005
2005
Earlier work this paper cites.
A. E. Gaweda, M. K. Muezzinoglu, G. R. Aronoff, A. A. Jacobs, J. M. Zurada, and M. E. Brier, “Reinforcement learning approach to individualization of chronic pharmacotherapy,” in IJCNN’05 , vol. 5. IEEE, 2005, pp. 3290–3295
2005
Earlier work this paper cites.
A. E. Gaweda, M. K. Muezzinoglu, G. R. Aronoff, A. A. Jacobs, J. M. Zurada, and M. E. Brier, “Individualization of pharmacological anemia management using reinforcement learning,” Neural Networks , vol. 18, no. 5-6, pp. 826–834, 2005
2005
Earlier work this paper cites.
E. D. Sinzinger and B. Moore, “Sedation of simulated icu patients using reinforcement learning based control,” International Journal on Artificial Intelligence Tools , vol. 14, no. 01n02, pp. 137–156, 2005
2005
Earlier work this paper cites.
J. Woodward, Making things happen: A theory of causal explanation . Oxford university press, 2005
2005
Earlier work this paper cites.
A. E. Gaweda, M. K. Muezzinoglu, G. R. Aronoff, A. A. Jacobs, J. M. Zurada, and M. E. Brier, “Incorporating prior knowledge into q-learning for drug delivery individualization,” in Fourth International Conference on Machine Learning and Applications . IEEE, 2005, pp. 6–pp
2005
Earlier work this paper cites.
A. E. Gaweda, M. K. Muezzinoglu, A. A. Jacobs, G. R. Aronoff, and M. E. Brier, “Model predictive control with reinforcement learning for drug delivery in renal anemia management,” in IEEE EMBS’06 . IEEE, 2006, pp. 5177–5180
2006
Earlier work this paper cites.
D. Ernst, G.-B. Stan, J. Goncalves, and L. Wehenkel, “Clinical data based optimal sti strategies for hiv: a reinforcement learning approach,” in 45th IEEE Conference on Decision and Control . IEEE, 2006, pp. 667–672
2006
Earlier work this paper cites.
N. Sadati, A. Aflaki, and M. Jahed, “Multivariable anesthesia control using reinforcement learning,” in IEEE SMC’06 , vol. 6. IEEE, 2006, pp. 4563–4568
2006
Earlier work this paper cites.
F. Sahba, H. R. Tizhoosh, and M. M. Salama, “A reinforcement learning framework for medical image segmentation,” in IJCNN , vol. 6, 2006, pp. 511–517
2006
Earlier work this paper cites.
S. J. Fakih and T. K. Das, “Lead: a methodology for learning efficient approaches to medical diagnosis,” IEEE Transactions on Information Technology in Biomedicine , vol. 10, no. 2, pp. 220–228, 2006
2006
Earlier work this paper cites.
D. Ramachandran and E. Amir, “Bayesian inverse reinforcement learning,” Urbana , vol. 51, no. 61801, pp. 1–4, 2007
2007
Earlier work this paper cites.
A. L. Strehl, C. Diuk, and M. L. Littman, “Efficient structure learning in factored-state mdps,” in AAAI , vol. 7, 2007, pp. 645–650
2007
Earlier work this paper cites.
S. A. Murphy, K. G. Lynch, D. Oslin, J. R. McKay, and T. TenHave, “Developing adaptive treatment strategies in substance abuse research,” Drug & Alcohol Dependence , vol. 88, pp. S24–S30, 2007
2007
Earlier work this paper cites.
P. Palumbo, S. Panunzi, and A. De Gaetano, “Qualitative behavior of a family of delay-differential models of the glucose-insulin system,” Discrete and Continuous Dynamical Systems Series B , vol. 7, no. 2, p. 399, 2007
2007
Earlier work this paper cites.
J. D. Martín-Guerrero, E. Soria-Olivas, M. Martínez-Sober, M. Climente-Martí, T. De Diego-Santos, and N. V. Jiménez-Torres, “Validation of a reinforcement learning policy for dosage optimization of erythropoietin,” in Australasian Joint Conference on Artificial Intelligence . Springer, 2007, pp. 732–738
2007
Earlier work this paper cites.
S. A. Murphy, D. W. Oslin, A. J. Rush, and J. Zhu, “Methodological challenges in constructing effective treatment sequences for chronic psychiatric disorders,” Neuropsychopharmacology , vol. 32, no. 2, p. 257, 2007
2007
Earlier work this paper cites.
J. Pineau, M. G. Bellemare, A. J. Rush, A. Ghizaru, and S. A. Murphy, “Constructing evidence-based treatment strategies using methods from computer science,” Drug & Alcohol Dependence , vol. 88, pp. S52–S60, 2007
2007
Earlier work this paper cites.
R. S. Keefe, R. M. Bilder, S. M. Davis, P. D. Harvey, B. W. Palmer, J. M. Gold, H. Y. Meltzer, M. F. Green, G. Capuano, T. S. Stroup et al. , “Neurocognitive effects of antipsychotic medications in patients with chronic schizophrenia in the catie trial,” Archives of General Psychiatry , vol. 64, no. 6, pp. 633–647, 2007
2007
Earlier work this paper cites.
M. Dennis and C. K. Scott, “Managing addiction as a chronic condition,” Addiction Science & Clinical Practice , vol. 4, no. 1, p. 45, 2007
2007
Earlier work this paper cites.
——, “Application of opposition-based reinforcement learning in image segmentation,” in 2007 IEEE Symposium on Computational Intelligence in Image and Signal Processing . IEEE, 2007, pp. 246–251
2007
Earlier work this paper cites.
J. Peters and S. Schaal, “Natural actor-critic,” Neurocomputing , vol. 71, no. 7-9, pp. 1180–1190, 2008
2008
Earlier work this paper cites.
L. Busoniu, R. Babuska, and B. De Schutter, “A comprehensive survey of multiagent reinforcement learning,” IEEE Transactions on Systems, Man, And Cybernetics-Part C: Applications and Reviews, 38 (2), 2008 , 2008
2008
Earlier work this paper cites.
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey, “Maximum entropy inverse reinforcement learning.” in AAAI , vol. 8. Chicago, IL, USA, 2008, pp. 1433–1438
2008
Earlier work this paper cites.
P. W. Lavori and R. Dawson, “Adaptive treatment strategies in chronic disease,” Annu. Rev. Med. , vol. 59, pp. 443–453, 2008
2008
Earlier work this paper cites.
A. Guez, R. D. Vincent, M. Avoli, and J. Pineau, “Adaptive treatment of epilepsy via batch-mode reinforcement learning.” in AAAI , 2008, pp. 1671–1678
2008
Earlier work this paper cites.
B. Chakraborty, V. Strecher, and S. Murphy, “Bias correction and confidence intervals for fitted q-iteration,” in Workshop on Model Uncertainty and Risk in Reinforcement Learning, NIPS, Whistler, Canada . Citeseer, 2008
2008
Earlier work this paper cites.
K. Krell, “Critical care workforce,” Critical Care Medicine , vol. 36, no. 4, pp. 1350–1353, 2008
2008
Earlier work this paper cites.
——, “Application of reinforcement learning for segmentation of transrectal ultrasound images,” BMC Medical Imaging , vol. 8, no. 1, p. 8, 2008
2008
Earlier work this paper cites.
S. M. B. Netto, V. R. C. Leite, A. C. Silva, A. C. de Paiva, and A. de Almeida Neto, “Application on reinforcement learning for diagnosis based on medical image,” in Reinforcement Learning . InTech, 2008
2008
Earlier work this paper cites.
K. Ragnarsson, “Functional electrical stimulation after spinal cord injury: current use, therapeutic effects and future directions,” Spinal cord , vol. 46, no. 4, p. 255, 2008
2008
Earlier work this paper cites.
P. S. Thomas, M. Branicky, A. Van Den Bogert, and K. Jagodnik, “Creating a reinforcement learning controller for functional electrical stimulation of a human arm,” in The Yale Workshop on Adaptive and Learning Systems , vol. 49326. NIH Public Access, 2008, p. 1
2008
Earlier work this paper cites.
V. L. Patel, E. H. Shortliffe, M. Stefanelli, P. Szolovits, M. R. Berthold, R. Bellazzi, and A. Abu-Hanna, “The coming of age of artificial intelligence in medicine,” Artificial Intelligence in Medicine , vol. 46, no. 1, pp. 5–17, 2009
2009
Earlier work this paper cites.
J. Kober and J. R. Peters, “Policy search for motor primitives in robotics,” in Advances in Neural Information Processing Systems , 2009, pp. 849–856
2009
Earlier work this paper cites.
A. L. Strehl, L. Li, and M. L. Littman, “Reinforcement learning in finite mdps: Pac analysis,” Journal of Machine Learning Research , vol. 10, no. Nov, pp. 2413–2444, 2009
2009
Earlier work this paper cites.
M. E. Taylor and P. Stone, “Transfer learning for reinforcement learning domains: A survey,” Journal of Machine Learning Research , vol. 10, no. Jul, pp. 1633–1685, 2009
2009
Earlier work this paper cites.
Y. Zhao, M. R. Kosorok, and D. Zeng, “Reinforcement learning design for cancer clinical trials,” Statistics in Medicine , vol. 28, no. 26, pp. 3294–3315, 2009
2009
Earlier work this paper cites.
S. Yasini, M. B. Naghibi Sistani, and A. Karimpour, “Agent-based simulation for blood glucose,” International Journal of Applied Science, Engineering and Technology , vol. 5, pp. 89–95, 2009
2009
Earlier work this paper cites.
B. P. Kovatchev, M. Breton, C. Dalla Man, and C. Cobelli, “In silico preclinical trials: a proof of concept in closed-loop control of type 1 diabetes,” 2009
2009
Earlier work this paper cites.
J. D. Martín-Guerrero, F. Gomez, E. Soria-Olivas, J. Schmidhuber, M. Climente-Martí, and N. V. Jiménez-Torres, “A reinforcement learning approach for individualizing erythropoietin dosages in hemodialysis patients,” Expert Systems with Applications , vol. 36, no. 6, pp. 9737–9742, 2009
2009
Earlier work this paper cites.
J. Pineau, A. Guez, R. Vincent, G. Panuccio, and M. Avoli, “Treating epilepsy via adaptive neurostimulation: a reinforcement learning approach,” International Journal of Neural Systems , vol. 19, no. 04, pp. 227–240, 2009
2009
Earlier work this paper cites.
K. Bush and J. Pineau, “Manifold embeddings for model-based reinforcement learning under partial observability,” in Advances in Neural Information Processing Systems , 2009, pp. 189–197
2009
Earlier work this paper cites.
R. S. Istepanian, N. Y. Philip, and M. G. Martini, “Medical qos provision based on reinforcement learning in ultrasound streaming over 3.5 g wireless systems,” IEEE Journal on Selected areas in Communications , vol. 27, no. 4, 2009
2009
Earlier work this paper cites.
P. S. Thomas, A. J. van den Bogert, K. M. Jagodnik, and M. S. Branicky, “Application of the actor-critic architecture to functional electrical stimulation control of a human arm.” in IAAI , 2009
2009
Earlier work this paper cites.
D. Koller, N. Friedman, and F. Bach, Probabilistic graphical models: principles and techniques . MIT press, 2009
2009
Earlier work this paper cites.
F. Elizalde, E. Sucar, J. Noguez, and A. Reyes, “Generating explanations based on markov decision processes,” in Mexican International Conference on Artificial Intelligence . Springer, 2009, pp. 51–62
2009
Earlier work this paper cites.
L. Busoniu, R. Babuska, B. De Schutter, and D. Ernst, Reinforcement learning and dynamic programming using function approximators . CRC press, 2010
2010
Earlier work this paper cites.
H. Xu and S. Mannor, “Distributionally robust markov decision processes,” in Advances in Neural Information Processing Systems , 2010, pp. 2505–2513
2010
Earlier work this paper cites.
A. Hassani et al. , “Reinforcement learning based control of tumor growth with chemotherapy,” in 2010 International Conference on System Science and Engineering (ICSSE) . IEEE, 2010, pp. 185–189
2010
Earlier work this paper cites.
M. Tenenbaum, A. Fern, L. Getoor, M. Littman, V. Manasinghka, S. Natarajan, D. Page, J. Shrager, Y. Singer, and P. Tadepalli, “Personalizing cancer therapy via machine learning,” in Workshops of NIPS , 2010
2010
Earlier work this paper cites.
E. Daskalaki, L. Scarnato, P. Diem, and S. G. Mougiakakou, “Preliminary results of a novel approach for glucose regulation using an actor-critic learning based controller,” 2010
2010
Earlier work this paper cites.
S. U. Acikgoz and U. M. Diwekar, “Blood glucose regulation with stochastic optimal control for insulin-dependent diabetic patients,” Chemical Engineering Science , vol. 65, no. 3, pp. 1227–1236, 2010
2010
Earlier work this paper cites.
A. Guez, “Adaptive control of epileptic seizures using reinforcement learning,” Ph.D. dissertation, McGill University Library, 2010
2010
Earlier work this paper cites.
B. Chakraborty, S. Murphy, and V. Strecher, “Inference for non-regular parameters in optimal dynamic treatment regimes,” Statistical Methods in Medical Research , vol. 19, no. 3, pp. 317–343, 2010
2010
Earlier work this paper cites.
B. L. Moore, P. Panousis, V. Kulkarni, L. D. Pyeatt, and A. G. Doufas, “Reinforcement learning for closed-loop propofol anesthesia: A human volunteer study.” in IAAI , 2010
2010
Earlier work this paper cites.
B. Zeng, A. Turkcan, J. Lin, and M. Lawley, “Clinic scheduling models with overbooking for patients with heterogeneous no-show probabilities,” Annals of Operations Research , vol. 178, no. 1, pp. 121–144, 2010
2010
Earlier work this paper cites.
D. J. Lizotte, M. H. Bowling, and S. A. Murphy, “Efficient reinforcement learning with multiple reward functions for randomized controlled trial analysis,” in ICML’10 . Citeseer, 2010, pp. 695–702
2010
Earlier work this paper cites.
S. Levine, Z. Popovic, and V. Koltun, “Nonlinear inverse reinforcement learning with gaussian processes,” in Advances in Neural Information Processing Systems , 2011, pp. 19–27
2011
Earlier work this paper cites.
I. Ahn and J. Park, “Drug scheduling of cancer chemotherapy based on natural actor-critic approach,” BioSystems , vol. 106, no. 2-3, pp. 121–129, 2011
2011
Earlier work this paper cites.
Y. Zhao, D. Zeng, M. A. Socinski, and M. R. Kosorok, “Reinforcement learning strategies for clinical trials in nonsmall cell lung cancer,” Biometrics , vol. 67, no. 4, pp. 1422–1433, 2011
2011
Earlier work this paper cites.
W. Cheng, J. Fürnkranz, E. Hüllermeier, and S.-H. Park, “Preference-based policy iteration: Leveraging preference learning for reinforcement learning,” in Joint European Conference on Machine Learning and Knowledge Discovery in Databases . Springer, 2011, pp. 312–327
2011
Earlier work this paper cites.
R. Eftimie, J. L. Bramson, and D. J. Earn, “Interactions between the immune system and cancer: a brief review of non-spatial mathematical models,” Bulletin of Mathematical Biology , vol. 73, no. 1, pp. 2–32, 2011
2011
Earlier work this paper cites.
J. de Lope, D. Maravall et al. , “Robust high performance reinforcement learning through weighted k-nearest neighbors,” Neurocomputing , vol. 74, no. 8, pp. 1251–1259, 2011
2011
Earlier work this paper cites.
C. Cobelli, E. Renard, and B. Kovatchev, “Artificial pancreas: past, present, future,” Diabetes , vol. 60, no. 11, pp. 2672–2682, 2011
2011
Earlier work this paper cites.
P. Escandell-Montero, J. M. Martínez-Martínez, J. D. Martín-Guerrero, E. Soria-Olivas, J. Vila-Francés, and R. Magdalena-Benedito, “Adaptive treatment of anemia on hemodialysis patients: A reinforcement learning approach,” in CIDM2011 . IEEE, 2011, pp. 44–49
2011
Earlier work this paper cites.
S. M. Shortreed, E. Laber, D. J. Lizotte, T. S. Stroup, J. Pineau, and S. A. Murphy, “Informing sequential clinical decision-making through reinforcement learning: an empirical study,” Machine Learning , vol. 84, no. 1-2, pp. 109–136, 2011
2011
Earlier work this paper cites.
B. L. Moore, A. G. Doufas, and L. D. Pyeatt, “Reinforcement learning: a novel method for optimal control of propofol-induced hypnosis,” Anesthesia & Analgesia , vol. 112, no. 2, pp. 360–367, 2011
2011
Earlier work this paper cites.
B. L. Moore, T. M. Quasny, and A. G. Doufas, “Reinforcement learning versus proportional–integral–derivative control of hypnosis in a simulated intraoperative patient,” Anesthesia & Analgesia , vol. 112, no. 2, pp. 350–359, 2011
2011
Earlier work this paper cites.
E. C. Borera, B. L. Moore, A. G. Doufas, and L. D. Pyeatt, “An adaptive neural network filter for improved patient state estimation in closed-loop anesthesia control,” in IEEE ICTAI’11 . IEEE, 2011, pp. 41–46
2011
Earlier work this paper cites.
Z. Huang, W. M. van der Aalst, X. Lu, and H. Duan, “Reinforcement learning based resource allocation in business process management,” Data & Knowledge Engineering , vol. 70, no. 1, pp. 127–145, 2011
2011
Earlier work this paper cites.
J. Fürnkranz and E. Hüllermeier, “Preference learning,” in Encyclopedia of Machine Learning . Springer, 2011, pp. 789–795
2011
Earlier work this paper cites.
N. Vlassis, M. Ghavamzadeh, S. Mannor, and P. Poupart, “Bayesian reinforcement learning,” in Reinforcement Learning . Springer, 2012, pp. 359–386
2012
Earlier work this paper cites.
L. Li, “Sample complexity bounds of exploration,” in Reinforcement Learning . Springer, 2012, pp. 175–204
2012
Earlier work this paper cites.
H. Van Hasselt, “Reinforcement learning in continuous state and action spaces,” in Reinforcement learning . Springer, 2012, pp. 207–251
2012
Earlier work this paper cites.
T. M. Moldovan and P. Abbeel, “Safe exploration in markov decision processes,” in Proceedings of the 29th International Coference on International Conference on Machine Learning . Omnipress, 2012, pp. 1451–1458
2012
Earlier work this paper cites.
M. Wiering and M. Van Otterlo, “Reinforcement learning,” Adaptation, learning, and optimization , vol. 12, 2012
2012
Earlier work this paper cites.
S. Lange, T. Gabel, and M. Riedmiller, “Batch reinforcement learning,” in Reinforcement learning . Springer, 2012, pp. 45–73
2012
Earlier work this paper cites.
T. Hester and P. Stone, “Learning and using models,” in Reinforcement learning . Springer, 2012, pp. 111–141
2012
Cited alongside, same era.
C. B. Browne, E. Powley, D. Whitehouse, S. M. Lucas, P. I. Cowling, P. Rohlfshagen, S. Tavener, D. Perez, S. Samothrakis, and S. Colton, “A survey of monte carlo tree search methods,” IEEE Transactions on Computational Intelligence and AI in games , vol. 4, no. 1, pp. 1–43, 2012
2012
Cited alongside, same era.
A. Lazaric, “Transfer in reinforcement learning: a framework and a survey,” in Reinforcement Learning . Springer, 2012, pp. 143–173
2012
Cited alongside, same era.
J. Fürnkranz, E. Hüllermeier, W. Cheng, and S.-H. Park, “Preference-based reinforcement learning: a formal framework and a policy iteration algorithm,” Machine Learning , vol. 89, no. 1-2, pp. 123–156, 2012
2012
Cited alongside, same era.
A. Jalalimanesh, H. S. Haghighi, A. Ahmadi, H. Hejazian, and M. Soltani, “Multi-objective optimization of radiotherapy: distributed q-learning and agent-based simulation,” Journal of Experimental & Theoretical Artificial Intelligence , pp. 1–16, 2017
2017
Later among the works it cites.
B. Stewart, C. P. Wild et al. , “World cancer report 2014,” Health , 2017
2017
Later among the works it cites.
A. Noori, M. A. Sadrnia et al. , “Glucose level control using temporal difference methods,” in 2017 Iranian Conference on Electrical Engineering (ICEE) . IEEE, 2017, pp. 895–900
2017
Later among the works it cites.
S. Parbhoo, J. Bogojeska, M. Zazzi, V. Roth, and F. Doshi-Velez, “Combining kernel and model based learning for hiv therapy selection,” AMIA Summits on Translational Science Proceedings , vol. 2017, p. 239, 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Zhifei and E. Meng Joo, “A survey of inverse reinforcement learning techniques,” International Journal of Intelligent Computing and Cybernetics , vol. 5, no. 3, pp. 293–311, 2012
2012
Cited alongside, same era.
B. Hengst, “Hierarchical approaches,” in Reinforcement learning . Springer, 2012, pp. 293–323
2012
Cited alongside, same era.
M. van Otterlo, “Solving relational and first-order logical markov decision processes: A survey,” in Reinforcement Learning . Springer, 2012, pp. 253–292
2012
Cited alongside, same era.
R. Akrour, M. Schoenauer, and M. Sebag, “April: Active preference learning-based reinforcement learning,” in Joint European Conference on Machine Learning and Knowledge Discovery in Databases . Springer, 2012, pp. 116–131
2012
Cited alongside, same era.
Y. Goldberg and M. R. Kosorok, “Q-learning with censored data,” Annals of Statistics , vol. 40, no. 1, p. 529, 2012
2012
Cited alongside, same era.
D. J. Lizotte, M. Bowling, and S. A. Murphy, “Linear fitted-q iteration with multiple reward functions,” Journal of Machine Learning Research , vol. 13, no. Nov, pp. 3253–3295, 2012
2012
Cited alongside, same era.
A. D. T. Force, V. Ranieri, G. Rubenfeld et al. , “Acute respiratory distress syndrome,” Jama , vol. 307, no. 23, pp. 2526–2533, 2012
2012
Cited alongside, same era.
W. M. Haddad, J. M. Bailey, B. Gholami, and A. R. Tannenbaum, “Clinical decision support and closed-loop control for intensive care unit sedation,” Asian Journal of Control , vol. 20, no. 5, pp. 1343–1350, 2012
2012
Cited alongside, same era.
T. W. Killian, S. Daulton, G. Konidaris, and F. Doshi-Velez, “Robust and efficient transfer learning with hidden parameter markov decision processes,” in Advances in Neural Information Processing Systems , 2017, pp. 6250–6261
2017
Later among the works it cites.
V. Nagaraj, A. Lamperski, and T. I. Netoff, “Seizure control in a computational model using a reinforcement learning stimulation paradigm,” International Journal of Neural Systems , vol. 27, no. 07, p. 1750012, 2017
2017
Later among the works it cites.
——, “Interactive q-learning for quantiles,” Journal of the American Statistical Association , vol. 112, no. 518, pp. 638–649, 2017
2017
Later among the works it cites.
E. L. Butler, E. B. Laber, S. M. Davis, and M. R. Kosorok, “Incorporating patient preferences into estimation of optimal individualized treatment rules,” Biometrics , 2017
2017
Later among the works it cites.
A. Rhodes, L. E. Evans, W. Alhazzani, M. M. Levy, M. Antonelli, R. Ferrer, A. Kumar, J. E. Sevransky, C. L. Sprung, M. E. Nunnally et al. , “Surviving sepsis campaign: international guidelines for management of sepsis and septic shock: 2016,” Intensive Care Medicine , vol. 43, no. 3, pp. 304–377, 2017
2017
Later among the works it cites.
T. Kamio, T. Van, and K. Masamune, “Use of machine-learning approaches to predict clinical deterioration in critically ill patients: A systematic review,” International Journal of Medical Research and Health Sciences , vol. 6, no. 6, pp. 1–7, 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
A. Raghu, M. Komorowski, L. A. Celi, P. Szolovits, and M. Ghassemi, “Continuous state-space models for optimal sepsis treatment: a deep reinforcement learning approach,” in Machine Learning for Healthcare Conference , 2017, pp. 147–163
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
E. F. Krakow, M. Hemmer, T. Wang, B. Logan, M. Arora, S. Spellman, D. Couriel, A. Alousi, J. Pidala, M. Last et al. , “Tools for the precision medicine era: How to develop highly personalized treatment recommendations from cohort and registry data using q-learning,” American journal of epidemiology , vol. 186, no. 2, pp. 160–172, 2017
2017
Later among the works it cites.
Y. Liu, B. Logan, N. Liu, Z. Xu, J. Tang, and Y. Wang, “Deep reinforcement learning for dynamic treatment regimes on medical registry data,” in IEEE ICHI’17 . IEEE, 2017, pp. 380–385
2017
Later among the works it cites.
2017
Later among the works it cites.
S. Jaber, G. Bellani, L. Blanch, A. Demoule, A. Esteban, L. Gattinoni, C. Guérin, N. Hill, J. G. Laffey, S. M. Maggiore et al. , “The intensive care medicine research agenda for airways, invasive and noninvasive mechanical ventilation,” Intensive Care Medicine , vol. 43, no. 9, pp. 1352–1365, 2017
2017
Later among the works it cites.
A. De Jong, G. Citerio, and S. Jaber, “Focus on ventilation and airway management in the icu,” Intensive Care Medicine , vol. 43, no. 12, pp. 1912–1915, 2017
2017
Later among the works it cites.
M. Fatima and M. Pasha, “Survey of machine learning algorithms for disease diagnostic,” Journal of Intelligent Learning Systems and Applications , vol. 9, no. 01, p. 1, 2017
2017
Later among the works it cites.
K. T. Chui, W. Alhalabi, S. S. H. Pang, P. O. d. Pablos, R. W. Liu, and M. Zhao, “Disease diagnosis in smart healthcare: Innovation, technologies and applications,” Sustainability , vol. 9, no. 12, p. 2309, 2017
2017
Later among the works it cites.
Y. Ling, S. A. Hasan, V. Datla, A. Qadir, K. Lee, J. Liu, and O. Farri, “Diagnostic inferencing via improving clinical concept extraction with deep reinforcement learning: A preliminary study,” in Machine Learning for Healthcare Conference , 2017, pp. 271–285
2017
Later among the works it cites.
F. C. Ghesu, B. Georgescu, Y. Zheng, S. Grbic, A. Maier, J. Hornegger, and D. Comaniciu, “Multi-scale deep reinforcement learning for real-time 3d-landmark detection in ct scans,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2017
2017
Later among the works it cites.
R. Liao, S. Miao, P. de Tournemire, S. Grbic, A. Kamen, T. Mansi, and D. Comaniciu, “An artificial agent for robust image registration.” in AAAI , 2017, pp. 4168–4175
2017
Later among the works it cites.
K. Ma, J. Wang, V. Singh, B. Tamersoy, Y.-J. Chang, A. Wimmer, and T. Chen, “Multimodal image registration with deep context reinforcement learning,” in International Conference on Medical Image Computing and Computer-Assisted Intervention . Springer, 2017, pp. 240–248
2017
Later among the works it cites.
J. Krebs, T. Mansi, H. Delingette, L. Zhang, F. C. Ghesu, S. Miao, A. K. Maier, N. Ayache, R. Liao, and A. Kamen, “Robust non-rigid registration through agent-based action learning,” in International Conference on Medical Image Computing and Computer-Assisted Intervention . Springer, 2017, pp. 344–352
2017
Later among the works it cites.
G. Maicas, G. Carneiro, A. P. Bradley, J. C. Nascimento, and I. Reid, “Deep reinforcement learning for active breast lesion detection from dce-mri,” in International Conference on Medical Image Computing and Computer-Assisted Intervention . Springer, 2017, pp. 665–673
2017
Later among the works it cites.
Y. Ling, S. A. Hasan, V. Datla, A. Qadir, K. Lee, J. Liu, and O. Farri, “Learning to diagnose: Assimilating clinical narratives using deep reinforcement learning,” in Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , vol. 1, 2017, pp. 895–905
2017
Later among the works it cites.
E. Y. Chang, M.-H. Wu, K.-F. T. Tang, H.-C. Kao, and C.-N. Chou, “Artificial intelligence in xprize deepq tricorder,” in Proceedings of the 2nd International Workshop on Multimedia for Personal Health and Health Care . ACM, 2017, pp. 11–18
2017
Later among the works it cites.
E. Y. Chang, “Deepq: Advancing healthcare through artificial intelligence and virtual reality,” in Proceedings of the 2017 ACM on Multimedia Conference . ACM, 2017, pp. 1068–1068
2017
Later among the works it cites.
T. S. M. T. Gomes, “Reinforcement learning for primary care e appointment scheduling,” 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
B. Thananjeyan, A. Garg, S. Krishnan, C. Chen, L. Miller, and K. Goldberg, “Multilateral surgical pattern cutting in 2d orthotropic gauze with deep reinforcement learning policies for tensioning,” in IEEE ICRA’17 . IEEE, 2017, pp. 2371–2378
2017
Later among the works it cites.
K. M. Jagodnik, P. S. Thomas, A. J. van den Bogert, M. S. Branicky, and R. F. Kirsch, “Training an actor-critic reinforcement learning controller for arm movement using human-generated rewards,” IEEE Transactions on Neural Systems and Rehabilitation Engineering , vol. 25, no. 10, pp. 1892–1905, 2017
2017
Later among the works it cites.
M. Olivecrona, T. Blaschke, O. Engkvist, and H. Chen, “Molecular de-novo design through deep reinforcement learning,” Journal of Cheminformatics , vol. 9, no. 1, p. 48, 2017
2017
Later among the works it cites.
E. Yom-Tov, G. Feraru, M. Kozdoba, S. Mannor, M. Tennenholtz, and I. Hochberg, “Encouraging physical activity in patients with diabetes: Intervention using a reinforcement learning system,” Journal of Medical Internet Research , vol. 19, no. 10, 2017
2017
Later among the works it cites.
A. Baniya, S. Herrmann, Q. Qiao, and H. Lu, “Adaptive interventions treatment modelling and regimen optimization using sequential multiple assignment randomized trials (smart) and q-learning,” in Proceedings of IIE Annual Conference , 2017, pp. 1187–1192
2017
Later among the works it cites.
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. P. Abbeel, and W. Zaremba, “Hindsight experience replay,” in NIPS’17 , 2017, pp. 5048–5058
2017
Later among the works it cites.
S. Racanière, T. Weber, D. Reichert, L. Buesing, A. Guez, D. J. Rezende, A. P. Badia, O. Vinyals, N. Heess, Y. Li et al. , “Imagination-augmented agents for deep reinforcement learning,” in Advances in neural information processing systems , 2017, pp. 5690–5701
2017
Later among the works it cites.
J. Fu, J. Co-Reyes, and S. Levine, “Ex2: Exploration with exemplar models for deep reinforcement learning,” in Advances in Neural Information Processing Systems , 2017, pp. 2574–2584
2017
Later among the works it cites.
H. Tang, R. Houthooft, D. Foote, A. Stooke, O. X. Chen, Y. Duan, J. Schulman, F. DeTurck, and P. Abbeel, “# exploration: A study of count-based exploration for deep reinforcement learning,” in NIPS’17 , 2017, pp. 2750–2759
2017
Later among the works it cites.
Z. C. Lipton, “The doctor just won’t accept that!” arXiv preprint arXiv:1711.08037 , 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
S. W. Carden and J. Livsey, “Small-sample reinforcement learning: Improving policies using synthetic data 1,” Intelligent Decision Technologies , vol. 11, no. 2, pp. 167–175, 2017
2017
Later among the works it cites.
J. Salamon and J. P. Bello, “Deep convolutional neural networks and data augmentation for environmental sound classification,” IEEE Signal Processing Letters , vol. 24, no. 3, pp. 279–283, 2017
2017
Later among the works it cites.
B. M. Lake, T. D. Ullman, J. B. Tenenbaum, and S. J. Gershman, “Building machines that learn and think like people,” Behavioral and Brain Sciences , vol. 40, 2017
2017
Later among the works it cites.
T. Ching, D. S. Himmelstein, B. K. Beaulieu-Jones, A. A. Kalinin, B. T. Do, G. P. Way, E. Ferrero, P.-M. Agapow, M. Zietz, M. M. Hoffman et al. , “Opportunities and obstacles for deep learning in biology and medicine,” bioRxiv , p. 142760, 2018
2018
Later among the works it cites.
Y. Li, “Deep reinforcement learning,” arXiv preprint arXiv:1810.06339 , 2018
2018
Later among the works it cites.
M. Mahmud, M. S. Kaiser, A. Hussain, and S. Vassanelli, “Applications of deep learning and reinforcement learning to biological data,” IEEE transactions on neural networks and learning systems , vol. 29, no. 6, pp. 2063–2079, 2018
2018
Later among the works it cites.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction . MIT press, 2018
2018
Later among the works it cites.
L. Buşoniu, T. de Bruin, D. Tolić, J. Kober, and I. Palunko, “Reinforcement learning for control: Performance, stability, and deep approximators,” Annual Reviews in Control , 2018
2018
Later among the works it cites.
D. Hein, S. Udluft, and T. A. Runkler, “Interpretable policies for reinforcement learning by genetic programming,” Engineering Applications of Artificial Intelligence , vol. 76, pp. 158–169, 2018
2018
Later among the works it cites.
O. Bastani, Y. Pu, and A. Solar-Lezama, “Verifiable reinforcement learning via policy extraction,” in Advances in Neural Information Processing Systems , 2018, pp. 2499–2509
2018
Later among the works it cites.
G. Yauney and P. Shah, “Reinforcement learning with action-derived rewards for chemotherapy and clinical trial dosing regimen selection,” in Machine Learning for Healthcare Conference , 2018, pp. 161–226
2018
Later among the works it cites.
M. Feng, G. Valdes, N. Dixit, and T. D. Solberg, “Machine learning in radiation oncology: Opportunities, requirements, and needs,” Frontiers in Oncology , vol. 8, 2018
2018
Later among the works it cites.
N. Cho, J. Shaw, S. Karuranga, Y. Huang, J. da Rocha Fernandes, A. Ohlrogge, and B. Malanda, “Idf diabetes atlas: Global estimates of diabetes prevalence for 2017 and projections for 2045,” Diabetes Research and Clinical Practice , vol. 138, pp. 271–281, 2018
2018
Later among the works it cites.
Q. Sun, M. Jankovic, J. Budzinski, B. Moore, P. Diem, C. Stettler, and S. G. Mougiakakou, “A dual mode adaptive basal-bolus advisor based on reinforcement learning,” IEEE journal of biomedical and health informatics , 2018
2018
Later among the works it cites.
P. D. Ngo, S. Wei, A. Holubová, J. Muzik, and F. Godtliebsen, “Reinforcement-learning optimal control for type-1 diabetes,” in 2018 IEEE EMBS International Conference on Biomedical & Health Informatics (BHI) . IEEE, 2018, pp. 333–336
2018
Later among the works it cites.
——, “Control of blood glucose for type-1 diabetes by using reinforcement learning with feedforward algorithm,” Computational and Mathematical Methods in Medicine , vol. 2018, 2018
2018
Later among the works it cites.
D. J. Luckett, E. B. Laber, A. R. Kahkoska, D. M. Maahs, E. Mayer-Davis, and M. R. Kosorok, “Estimating dynamic treatment regimes in mobile health using v-learning,” Journal of the American Statistical Association , no. just-accepted, pp. 1–39, 2018
2018
Later among the works it cites.
J. Yao, T. Killian, G. Konidaris, and F. Doshi-Velez, “Direct policy transfer via hidden parameter markov decision processes,” 2018
2018
Later among the works it cites.
G. Panuccio, M. Semprini, L. Natale, S. Buccelli, I. Colombi, and M. Chiappalone, “Progress in neuroengineering for brain repair: New challenges and open issues,” Brain and Neuroscience Advances , vol. 2, p. 2398212818776475, 2018
2018
Later among the works it cites.
Y. Tao, L. Wang, D. Almirall et al. , “Tree-based reinforcement learning for estimating optimal dynamic treatment regimes,” The Annals of Applied Statistics , vol. 12, no. 3, pp. 1914–1938, 2018
2018
Later among the works it cites.
A. Vellido, V. Ribas, C. Morales, A. R. Sanmartín, and J. C. R. Rodríguez, “Machine learning in critical care: state-of-the-art and a sepsis case study,” Biomedical engineering online , vol. 17, no. 1, p. 135, 2018
2018
Later among the works it cites.
M. Komorowski, L. A. Celi, O. Badawi, A. C. Gordon, and A. A. Faisal, “The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care,” Nature Medicine , vol. 24, no. 11, p. 1716, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
C. P. Utomo, X. Li, and W. Chen, “Treatment recommendation in critical care: A scalable and interpretable approach in partially observable health states,” 2018
2018
Later among the works it cites.
J. Futoma, A. Lin, M. Sendak, A. Bedoya, M. Clement, C. O’Brien, and K. Heller, “Learning to treat sepsis with multi-output gaussian process deep recurrent q-networks,” 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
R. Lin, M. D. Stanley, M. M. Ghassemi, and S. Nemati, “A deep deterministic policy gradient approach to medication dosing and surveillance in the icu,” in IEEE EMBC’18 . IEEE, 2018, pp. 4927–4931
2018
Later among the works it cites.
L. Wang, W. Zhang, X. He, and H. Zha, “Supervised reinforcement learning with recurrent neural network for dynamic treatment recommendation,” in Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . ACM, 2018, pp. 2447–2456
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
S. Saria, “Individualized sepsis treatment using reinforcement learning,” Nature medicine , vol. 24, no. 11, p. 1641, 2018
2018
Later among the works it cites.
S. K. Rai and K. Sowmya, “A review on use of machine learning techniques in diagnostic health-care,” Artificial Intelligent Systems and Machine Learning , vol. 10, no. 4, pp. 102–107, 2018
2018
Later among the works it cites.
A. Bernstein and E. Burnaev, “Reinforcement learning in computer vision,” in CMV’17 , vol. 10696. International Society for Optics and Photonics, 2018, p. 106961S
2018
Later among the works it cites.
D. Liu and T. Jiang, “Deep reinforcement learning for surgical gesture segmentation and classification,” in International Conference on Medical Image Computing and Computer-Assisted Intervention . Springer, 2018, pp. 247–255
2018
Later among the works it cites.
F. C. Ghesu, B. Georgescu, S. Grbic, A. Maier, J. Hornegger, and D. Comaniciu, “Towards intelligent robust detection of anatomical structures in incomplete volumetric data,” Medical Image Analysis , vol. 48, pp. 203–213, 2018
2018
Later among the works it cites.
M. Etcheverry, B. Georgescu, B. Odry, T. J. Re, S. Kaushik, B. Geiger, N. Mariappan, S. Grbic, and D. Comaniciu, “Nonlinear adaptively learned optimization for object localization in 3d medical images,” in Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support . Springer, 2018, pp. 254–262
2018
Later among the works it cites.
A. Alansary, O. Oktay, Y. Li, L. Le Folgoc, B. Hou, G. Vaillant, B. Glocker, B. Kainz, and D. Rueckert, “Evaluating reinforcement learning agents for anatomical landmark detection,” 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
P. Zhang, F. Wang, and Y. Zheng, “Deep reinforcement learning for vessel centerline tracing in multi-modality 3d volumes,” in International Conference on Medical Image Computing and Computer-Assisted Intervention . Springer, 2018, pp. 755–763
2018
Later among the works it cites.
H.-C. Kao, K.-F. Tang, and E. Y. Chang, “Context-aware symptom checking for disease diagnosis using hierarchical reinforcement learning,” 2018
2018
Later among the works it cites.
Z. Wei, Q. Liu, B. Peng, H. Tou, T. Chen, X. Huang, K.-F. Wong, and X. Dai, “Task-oriented dialogue system for automatic diagnosis,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , vol. 2, 2018, pp. 201–207
2018
Later among the works it cites.
2018
Later among the works it cites.
D. Baek, M. Hwang, H. Kim, and D.-S. Kwon, “Path planning for automation of surgery robot based on probabilistic roadmap and reinforcement learning,” in 2018 15th International Conference on Ubiquitous Robots (UR) . IEEE, 2018, pp. 342–347
2018
Later among the works it cites.
K. Li, M. Rath, and J. W. Burdick, “Inverse reinforcement learning via function approximation for clinical motion analysis,” in IEEE ICRA’18 . IEEE, 2018, pp. 610–617
2018
Later among the works it cites.
A. Serrano, B. Imbernón, H. Pérez-Sánchez, J. M. Cecilia, A. Bueno-Crespo, and J. L. Abellán, “Accelerating drugs discovery with deep reinforcement learning: An early approach,” in Proceedings of the 47th International Conference on Parallel Processing Companion . ACM, 2018, p. 6
2018
Later among the works it cites.
D. Neil, M. Segler, L. Guasch, M. Ahmed, D. Plumbley, M. Sellwood, and N. Brown, “Exploring deep recurrent models with reinforcement learning for molecule design,” 2018
2018
Later among the works it cites.
M. Popova, O. Isayev, and A. Tropsha, “Deep reinforcement learning for de novo drug design,” Science Advances , vol. 4, no. 7, p. eaap7885, 2018
2018
Later among the works it cites.
E. M. Forman, S. G. Kerrigan, M. L. Butryn, A. S. Juarascio, S. M. Manasse, S. Ontañón, D. H. Dallal, R. J. Crochiere, and D. Moskow, “Can the artificial intelligence technique of reinforcement learning use continuously-monitored digital data to optimize treatment for weight loss?” Journal of Behavioral Medicine , pp. 1–15, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
M. Dimakopoulou and B. Van Roy, “Coordinated exploration in concurrent reinforcement learning,” in International Conference on Machine Learning , 2018, pp. 1270–1278
2018
Later among the works it cites.
T. Mannucci, E.-J. van Kampen, C. de Visser, and Q. Chu, “Safe exploration algorithms for reinforcement learning controllers,” IEEE transactions on neural networks and learning systems , vol. 29, no. 4, pp. 1069–1081, 2018
2018
Later among the works it cites.
Z. C. Lipton, “The mythos of model interpretability,” Communications of the ACM , vol. 61, no. 10, pp. 36–43, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
A. Verma, V. Murali, R. Singh, P. Kohli, and S. Chaudhuri, “Programmatically interpretable reinforcement learning,” in International Conference on Machine Learning , 2018, pp. 5052–5061
2018
Later among the works it cites.
M. Wu, M. C. Hughes, S. Parbhoo, M. Zazzi, V. Roth, and F. Doshi-Velez, “Beyond sparsity: Tree regularization of deep models for interpretability,” in Thirty-Second AAAI Conference on Artificial Intelligence , 2018
2018
Later among the works it cites.
C. Yu, D. Wang, T. Yang, W. Zhu, Y. Li, H. Ge, and J. Ren, “Adaptively shaping reinforcement learning agents via human reward,” in Pacific Rim International Conference on Artificial Intelligence . Springer, 2018, pp. 85–97
2018
Later among the works it cites.
2018
Later among the works it cites.
F. Zhu, J. Guo, R. Li, and J. Huang, “Robust actor-critic contextual bandit for mobile health (mhealth) interventions,” in Proceedings of the 2018 ACM International Conference on Bioinformatics, Computational Biology, and Health Informatics . ACM, 2018, pp. 492–501
2018
Later among the works it cites.
F. Zhu, J. Guo, Z. Xu, P. Liao, L. Yang, and J. Huang, “Group-driven reinforcement learning for personalized mhealth intervention,” in International Conference on Medical Image Computing and Computer-Assisted Intervention . Springer, 2018, pp. 590–598
2018
Later among the works it cites.
J. He, S. L. Baxter, J. Xu, J. Xu, X. Zhou, and K. Zhang, “The practical implementation of artificial intelligence technologies in medicine,” Nature Medicine , vol. 25, no. 1, p. 30, 2019
2019
Closest in time.
A. Esteva, A. Robicquet, B. Ramsundar, V. Kuleshov, M. DePristo, K. Chou, C. Cui, G. Corrado, S. Thrun, and J. Dean, “A guide to deep learning in healthcare,” Nature Medicine , vol. 25, no. 1, p. 24, 2019
2019
Closest in time.
O. Gottesman, F. Johansson, M. Komorowski, A. Faisal, D. Sontag, F. Doshi-Velez, and L. A. Celi, “Guidelines for reinforcement learning in healthcare.” Nature medicine , vol. 25, no. 1, p. 16, 2019
2019
Closest in time.
C. Yu, Y. Dong, J. Liu, and G. Ren, “Incorporating causal factors into reinforcement learning for dynamic treatment regimes in hiv,” BMC medical informatics and decision making , vol. 19, no. 2, p. 60, 2019
2019
Closest in time.
2019
Closest in time.
C. Yu, G. Ren, and J. Liu, “Deep inverse reinforcement learning for sepsis treatment,” in 2019 IEEE ICHI , 2019, pp. 1–3
2019
Closest in time.
C. Yu, J. Liu, and H. Zhao, “Inverse reinforcement learning for intelligent mechanical ventilation and sedative dosing in intensive care units,” BMC medical informatics and decision making , vol. 19, no. 2, p. 57, 2019
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
E. J. Topol, “High-performance medicine: the convergence of human and artificial intelligence,” Nature medicine , vol. 25, no. 1, p. 44, 2019
2019
Closest in time.
C. Yu, G. Ren, and Y. Dong, “Supervised-actor-critic reinforcement learning for intelligent mechanical ventilation and sedative dosing in intensive care units,” BMC medical informatics and decision making , 2020
2020
Closest in time.
J. M. Malof and A. E. Gaweda, “Optimizing drug therapy with reinforcement learning: The case of anemia management,” in Neural Networks (IJCNN), The 2011 International Joint Conference on . IEEE, 2011, pp. 2088–2092
2092
Closest in time.