Fetching the paper…
Reading the bibliography…
Hallucinations in foundation models arise from autoregressive training objectives that prioritize token-likelihood optimization over epistemic accuracy, fostering overconfidence and poorly calibrated uncertainty.
Jaccard, P.: Étude comparative de la distribution florale dans une portion des alpes et des jura. Bull Soc Vaudoise Sci Nat 37
1901
Earlier work this paper cites.
Tversky, A., Kahneman, D.: Judgment under uncertainty: Heuristics and biases: Biases in judgments reveal some heuristics of thinking under uncertainty. Science 185
1974
Earlier work this paper cites.
Shekelle, P.G., Ortiz, E., Rhodes, S., Morton, S.C., Eccles, M.P., Grimshaw, J.M., Woolf, S.H.: Developing clinical guidelines. Western Journal of Medicine 176
2002
Earlier work this paper cites.
Finch, A., Hwang, Y.-S., Sumita, E.: Using machine translation evaluation techniques to determine sentence-level semantic equivalence. In: Proceedings of the Third International Workshop on Paraphrasing (IWP2005) (2005)
2005
Earlier work this paper cites.
Donnelly, K., et al
2006
Earlier work this paper cites.
Pradhan, S., Elhadad, N., Chapman, W., Manandhar, S., Savova, G.: SemEval-2014 task 7: Analysis of clinical text. In: Nakov, P., Zesch, T. (eds.) Proceedings of the 8th International Workshop on Semantic Evaluation (SemEval 2014), pp. 54–62. Association for Computational Linguistics, Dublin, Ireland (2014). https://doi.org/10.3115/v1/S14-2007 . https://aclanthology.org/S14-2007/
2007
Earlier work this paper cites.
Padó, S., Cer, D., Galley, M., Jurafsky, D., Manning, C.D.: Measuring machine translation quality as semantic equivalence: A metric based on entailment features. Machine Translation 23
2009
Earlier work this paper cites.
O’Brien, D.T.: Thinking, fast and slow by daniel kahneman. (2012)
2012
Earlier work this paper cites.
Xia, F., Yetisgen-Yildiz, M.: Clinical corpus annotation: challenges and strategies. In: Proceedings of the Third Workshop on Building and Evaluating Resources for Biomedical Text Mining (BioTxtM’2012) in Conjunction with the International Conference on Language Resources and Evaluation (LREC), Istanbul, Turkey, pp. 21–27 (2012)
2012
Earlier work this paper cites.
Borden, N., Linklater, D.: Hickam’s dictum. Western Journal of Emergency Medicine: Integrating Emergency Care with Population Health 14
2013
Earlier work this paper cites.
Blumenthal-Barby, J.S., Krieger, H.: Cognitive biases and heuristics in medical decision making: A critical review using a systematic search strategy. Medical Decision Making 35
2014
Earlier work this paper cites.
Karimi, S., Metke-Jimenez, A., Kemp, M., Wang, C.: Cadec: A corpus of adverse drug event annotations. Journal of Biomedical Informatics 55
2015
Earlier work this paper cites.
Svenstrup, D., Jørgensen, H.L., Winther, O.: Rare disease diagnosis: a review of web search, social media and large-scale data-mining approaches. Rare Diseases 3
2015
Earlier work this paper cites.
Bari, A., Khan, R.A., Rathore, A.W.: Medical errors; causes, consequences, emotional response and resulting behavioral change. Pakistan Journal of Medical Sciences 32
2016
Earlier work this paper cites.
Saposnik, G., Redelmeier, D., Ruff, C.C., Tobler, P.N.: Cognitive biases associated with medical decisions: A systematic review. BMC Medical Informatics and Decision Making 16
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
Tsiaras, S.V., Safi, L.M., Ghoshhajra, B.B., Lindsay, M.E., Wood, M.J.: Case 39-2017: A 41-year-old woman with recurrent chest pain. New England Journal of Medicine 377
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Camburu, O.-M., Rocktäschel, T., Lukasiewicz, T., Blunsom, P.: e-snli: Natural language inference with natural language explanations. Advances in Neural Information Processing Systems 31
2018
Earlier work this paper cites.
Mehta, N., Devarakonda, M.V.: Machine learning, natural language programming, and electronic health records: The next step in the artificial intelligence journey? Journal of Allergy and Clinical Immunology 141
2018
Earlier work this paper cites.
Murphy, J.E., Shampain, K., Riley, L.E., Clark, J.W., Basnet, K.M.: Case 32-2018: A 36-year-old pregnant woman with newly diagnosed adenocarcinoma. New England Journal of Medicine 379
2018
Earlier work this paper cites.
AMA Journal of Ethics: Are current tort liability doctrines adequate for addressing injury caused by AI? AMA Journal of Ethics 21
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Chen, I.Y., Szolovits, P., Ghassemi, M.: Can ai help reduce disparities in general medical and mental health care? AMA journal of ethics 21
2019
Earlier work this paper cites.
Devlin, J., Chang, M.-W., Lee, K., Toutanova, K.: BERT: Pre-training of deep bidirectional transformers for language understanding. In: Burstein, J., Doran, C., Solorio, T. (eds.) Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4171–4186. Association for Computational Linguistics, Minneapolis, Minnesota (2019). https://doi.org/10.18653/v1/N19-1423 . https://aclanthology.org/N19-1423
2019
Earlier work this paper cites.
Falke, T., Ribeiro, L.F., Utama, P.A., Dagan, I., Gurevych, I.: Ranking generated summaries by correctness: An interesting but challenging application for natural language inference. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 2214–2220 (2019)
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Scialom, T., Lamprier, S., Piwowarski, B., Staiano, J.: Answers unite! unsupervised metrics for reinforced summarization models. In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 3246–3256 (2019)
2019
Earlier work this paper cites.
Topol, E.J.: High-performance medicine: the convergence of human and artificial intelligence. Nature medicine 25
2019
Earlier work this paper cites.
Walley, A.Y., Wakeman, S.E., Eng, G.: Case 6-2019: A 29-year-old woman with nausea, vomiting, and diarrhea. New England Journal of Medicine 380
2019
Earlier work this paper cites.
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., Amodei, D.: Language models are few-shot learners 33
2020
Earlier work this paper cites.
Desai, S., Durrett, G.: Calibration of pre-trained transformers. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 295–302 (2020). https://doi.org/10.18653/v1/2020.emnlp-main.20 . Association for Computational Linguistics
2020
Earlier work this paper cites.
Hammond, M.E.H., Stehlik, J., Drakos, S.G., Kfoury, A.G.: Bias in medicine: Lessons learned and mitigation strategies. JACC: Basic to Translational Science 6
2020
Earlier work this paper cites.
Kamath, A., Jia, R., Liang, P.: Selective question answering under domain shift. In: Jurafsky, D., Chai, J., Schluter, N., Tetreault, J. (eds.) Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 5684–5696. Association for Computational Linguistics, Online (2020). https://doi.org/10.18653/v1/2020.acl-main.503 . https://aclanthology.org/2020.acl-main.503
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Kazi, D.S., Martin, L.M., Litmanovich, D., Pinto, D.S., Clerkin, K.J., Zimetbaum, P.J., Dudzinski, D.M.: Case 18-2020: a 73-year-old man with hypoxemic respiratory failure and cardiac dysfunction. New England Journal of Medicine 382
2020
Earlier work this paper cites.
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Wang, A., Cho, K., Lewis, M.: Asking and answering questions to evaluate the factual consistency of summaries. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 5008–5020 (2020)
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
De Cao, N., Aziz, W., Titov, I.: Editing factual knowledge in language models. In: EMNLP 2021-2021 Conference on Empirical Methods in Natural Language Processing, Proceedings, pp. 6491–6506 (2021)
2021
Earlier work this paper cites.
Gong, F., Wang, M., Wang, H., Wang, S., Liu, M.: Smr: medical knowledge graph embedding for safe medicine recommendation. Big Data Research 23
2021
Earlier work this paper cites.
Jiang, Z., Araki, J., Ding, H., Neubig, G.: How can we know when language models know? on the calibration of language models for question answering. Transactions of the Association for Computational Linguistics 9
2021
Earlier work this paper cites.
Jin, D., Pan, E., Oufattole, N., Weng, W.-H., Fang, H., Szolovits, P.: What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences 11
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Rehana, R.W., Huda, N.: A common heuristic in medicine: anchoring. Ann Med Health Sci Res 11
2021
Earlier work this paper cites.
Scialom, T., Dray, P.-A., Lamprier, S., Piwowarski, B., Staiano, J., Wang, A., Gallinari, P.: Questeval: Summarization asks for fact-based evaluation. In: 2021 Conference on Empirical Methods in Natural Language Processing, pp. 6594–6604 (2021). Association for Computational Linguistics
2021
Earlier work this paper cites.
Yogatama, D., Masson d’Autume, C., Kong, L.: Adaptive semiparametric language models. Transactions of the Association for Computational Linguistics 9
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., Van Den Driessche, G.B., Lespiau, J.-B., Damoc, B., Clark, A., et al
2022
Earlier work this paper cites.
De Nicola, A., Zgheib, R., Taglino, F.: Toward a knowledge graph for medical diagnosis: issues and usage scenarios, 129–142 (2022)
2022
Earlier work this paper cites.
Hata, T., Shima, H., Nitta, M., Ueda, E., Nishihara, M., Uchiyama, K., Katsumata, T., Neo, M.: The relationship between duration of general anesthesia and postoperative fall risk during hospital stay in orthopedic patients. Journal of Patient Safety 18
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Meng, K., Bau, D., Andonian, A., Belinkov, Y.: Locating and editing factual associations in gpt. Advances in Neural Information Processing Systems 35
2022
Earlier work this paper cites.
Mitchell, E., Lin, C., Bosselut, A., Manning, C.D., Finn, C.: Memory-based model editing at scale. In: International Conference on Machine Learning, pp. 15817–15831 (2022). PMLR
2022
Earlier work this paper cites.
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al
2022
Earlier work this paper cites.
Pal, A., Umapathi, L.K., Sankarasubbu, M.: Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. In: Conference on Health, Inference, and Learning, pp. 248–260 (2022). PMLR
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
The White House Office of Science and Technology Policy: Blueprint for an AI Bill of Rights. Accessed: 2024-11-04 (2022). https://www.whitehouse.gov/ostp/ai-bill-of-rights/
2022
Earlier work this paper cites.
Wang, S., Lin, M., Ghosal, T., Ding, Y., Peng, Y.: Knowledge graph applications in medical imaging analysis: a scoping review. Health data science 2022
2022
Earlier work this paper cites.
Whitehead, S., Petryk, S., Shakib, V., Gonzalez, J., Darrell, T., Rohrbach, A., Rohrbach, M.: Reliable visual question answering: Abstain rather than answer incorrectly. In: European Conference on Computer Vision, pp. 148–166 (2022). Springer
2022
Earlier work this paper cites.
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q.V., Zhou, D., et al
2022
Earlier work this paper cites.
Yu, G., Tabatabaei, M., Mezei, J., Zhong, Q., Chen, S., Li, Z., Li, J., Shu, L., Shu, Q.: Improving chronic disease management for children with knowledge graphs and artificial intelligence. Expert Systems with Applications 201
2022
Earlier work this paper cites.
Asai, A., Min, S., Zhong, Z., Chen, D.: Retrieval-based language models and applications. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 6: Tutorial Abstracts), pp. 41–46 (2023)
2023
Earlier work this paper cites.
Abu-Salih, B., Al-Qurishi, M., Alweshah, M., Al-Smadi, M., Alfayez, R., Saadeh, H.: Healthcare knowledge graph construction: A systematic review of the state-of-the-art, open issues, and opportunities. Journal of Big Data 10
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Bottomley, D., Thaldar, D.: Liability for harm caused by AI in healthcare: An overview of the core legal concepts. Frontiers in Pharmacology 14
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Chandak, P., Huang, K., Zitnik, M.: Building a knowledge graph to enable precision medicine. Scientific Data 10
2023
Earlier work this paper cites.
Chen, S., Kann, B.H., Foote, M.B., Aerts, H.J.W.L., Savova, G.K., Mak, R.H., Bitterman, D.S.: Use of artificial intelligence chatbots for cancer treatment information. JAMA Oncology 9
2023
Earlier work this paper cites.
Chen, S., Kann, B.H., Foote, M.B., Aerts, H.J., Savova, G.K., Mak, R.H., Bitterman, D.S.: The utility of chatgpt for cancer treatment information. medrxiv. Preprint posted March 16
2023
Earlier work this paper cites.
Congressional Research Service: Overview of Artificial Intelligence: The National AI Strategy and Major U.S. Policy Issues. Accessed: 2024-11-04 (2023). https://crsreports.congress.gov/product/pdf/R/R47843
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Fanta, G.B., Pretorius, L.: Sociotechnical factors of sustainable digital health systems: A system dynamics model. Sustainable Futures (2023) https://doi.org/10.1016/j.sftr.2023.100127
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Group, M.M.W.: Federated benchmarking of medical artificial intelligence with medperf. Nature Machine Intelligence (2023)
2023
Earlier work this paper cites.
Guerreiro, N.M., Voita, E., Martins, A.F.: Looking for a needle in a haystack: A comprehensive study of hallucinations in neural machine translation. In: Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pp. 1059–1075 (2023)
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Hagendorff, T., Fabi, S., Kosinski, M.: Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in chatgpt. Nature Computational Science (2023) https://doi.org/10.1038/s43588-023-00527-x . Brief Communication; Received: 17 February 2023; Accepted: 5 September 2023; Published online: 5 October 2023
2023
Earlier work this paper cites.
Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., et al.: A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems (2023)
2023
Earlier work this paper cites.
Izacard, G., Lewis, P., Lomeli, M., Hosseini, L., Petroni, F., Schick, T., Dwivedi-Yu, J., Joulin, A., Riedel, S., Grave, E.: Atlas: Few-shot learning with retrieval augmented language models. Journal of Machine Learning Research 24
2023
Earlier work this paper cites.
Jin, Q., Kim, W., Chen, Q., Comeau, D.C., Yeganova, L., Wilbur, W.J., Lu, Z.: Medcpt: Contrastive pre-trained transformers with large-scale pubmed search logs for zero-shot biomedical information retrieval. Bioinformatics 39
2023
Earlier work this paper cites.
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y.J., Madotto, A., Fung, P.: Survey of hallucination in natural language generation. ACM Computing Surveys 55
2023
Cited alongside, same era.
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y.J., Madotto, A., Fung, P.: Survey of hallucination in natural language generation. ACM Computing Surveys 55
2023
Cited alongside, same era.
Jiang, H., Liu, P., Shang, J., Wang, F.: Recent advances in natural language processing for clinical medicine: A survey. Artificial Intelligence in Medicine 128
2023
Cited alongside, same era.
Juhi, A., Pipil, N., Santra, S., Mondal, S., Behera, J.K., Mondal, H.: The capability of chatgpt in predicting and explaining common drug-drug interactions. Cureus 15
2023
Cited alongside, same era.
2024
Later among the works it cites.
Kim, H.-K.: The effects of artificial intelligence chatbots on women’s health: A systematic review and meta-analysis. Healthcare 12
2024
Later among the works it cites.
2024
Later among the works it cites.
Ke, Y., Yang, R., Lie, S.A., Lim, T.X.Y., Ning, Y., Li, I., Abdullah, H.R., Ting, D.S.W., Liu, N.: Mitigating cognitive biases in clinical decision-making through multi-agent conversations using large language models: Simulation study. Journal of Medical Internet Research 26
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Koopman, B., Zuccon, G.: Dr chatgpt tell me what i want to hear: How different prompts impact health answer correctness. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 15012–15022 (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Lytal, Reiter, Smith, Ivey, Fronrath: Can AI Be Liable for Malpractice? Blog post. Accessed: 2025-02-16 (2023). https://www.foryourrights.com/blog/ai-and-medical-malpractice/
2023
Cited alongside, same era.
Ly, D.P., Shekelle, P.G., Song, Z.: Evidence for anchoring bias during physician decision-making. JAMA internal medicine 183
2023
Cited alongside, same era.
Mohammadi, A., Etemad, B., Zhang, X., Li, Y., Bedwell, G.J., Sharaf, R., Kittilson, A., Melberg, M., Crain, C.R., Traunbauer, A.K., Wong, C., Fajnzylber, J., Worrall, D.P., Rosenthal, A., Jordan, H., Jilg, N., Kaseke, C., Giguel, F., Lian, X., Deo, R., Gillespie, E., Chishti, R., Abrha, S., Adams, T., Li, J.Z.: Viral and host mediators of non-suppressible hiv-1 viremia. Nature Medicine 29
2023
Cited alongside, same era.
2024
Later among the works it cites.
Li, S.S., Balachandran, V., Feng, S., Ilgen, J.S., Pierson, E., Koh, P.W., Tsvetkov, Y.: Mediq: Question-asking llms and a benchmark for reliable interactive clinical reasoning. In: The Thirty-eighth Annual Conference on Neural Information Processing Systems (2024)
2024
Later among the works it cites.
2024
Later among the works it cites.
Lee, J., Park, S., Shin, J., Cho, B.: Analyzing evaluation methods for large language models in the medical field: a scoping review. BMC Medical Informatics and Decision Making 24
2024
Later among the works it cites.
2024
Later among the works it cites.
Li, Y., Yang, C., Ettinger, A.: When hindsight is not 20/20: Testing limits on reflective thinking in large language models. In: Duh, K., Gomez, H., Bethard, S. (eds.) Findings of the Association for Computational Linguistics: NAACL 2024, pp. 3741–3753. Association for Computational Linguistics, Mexico City, Mexico (2024). https://doi.org/10.18653/v1/2024.findings-naacl.237 . https://aclanthology.org/2024.findings-naacl.237
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
McDermott, M., Dighe, A., Szolovits, P., Luo, Y., Baron, J.: Using machine learning to develop smart reflex testing protocols. Journal of the American Medical Informatics Association 31
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., et al.: Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems 36
2024
Later among the works it cites.
2024
Later among the works it cites.
National Institute of Standards and Technology: Artificial Intelligence Risk Management Framework. Technical Report NIST AI 100-1, U.S. Department of Commerce (2023). Accessed: 2024-11-04. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf
2024
Later among the works it cites.
Nazi, Z.A., Peng, W.: Large language models in healthcare and medical domain: A review. Informatics 11
2024
Later among the works it cites.
2024
Later among the works it cites.
Pressman, S.M., Borna, S., Gomez-Cabello, C.A., Haider, S.A., Haider, C.R., Forte, A.J.: Clinical and surgical applications of large language models: A systematic review. Journal of Clinical Medicine 13
2024
Later among the works it cites.
Penkov, S.: Mitigating hallucinations in large language models via semantic enrichment of prompts: Insights from biobert and ontological integration. In: Proceedings of the Sixth International Conference on Computational Linguistics in Bulgaria (CLIB 2024), pp. 272–276 (2024)
2024
Later among the works it cites.
Rodriguez, J.A., Alsentzer, E., Bates, D.W.: Leveraging large language models to foster equity in healthcare. Journal of the American Medical Informatics Association, 055 (2024)
2024
Later among the works it cites.
Reddy, S.: Generative ai in healthcare: an implementation science informed translational path on application, integration and governance. Implementation Science 19
2024
Later among the works it cites.
Rafailov, R., Sharma, A., Mitchell, E., Manning, C.D., Ermon, S., Finn, C.: Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems 36
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Savage, T., Wang, J., Gallo, R., Boukil, A., Patel, V., Safavi-Naini, S.A., Soroush, A., Chen, J.H.: Large language model uncertainty proxies: Discrimination and calibration for medical diagnosis and treatment. Journal of the American Medical Informatics Association 31
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Szolovits, P.: Large Language Models Seem Miraculous, but Science Abhors Miracles. Massachusetts Medical Society (2024)
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
U.S. Department of Health and Human Services: Breach Notification Rule. Accessed: 2024-11-04 (2024). https://www.hhs.gov/hipaa/for-professionals/breach-notification/index.html
2024
Later among the works it cites.
U.S. Department of Health and Human Services: Health Information Privacy. Accessed: 2024-11-04 (2024). https://www.hhs.gov/hipaa/for-professionals/privacy/index.html
2024
Later among the works it cites.
U.S. Department of Health and Human Services: Health Information Privacy: Laws & Regulations. https://www.hhs.gov/hipaa/for-professionals/privacy/laws-regulations/index.html . Accessed on November 4, 2024 (2024)
2024
Later among the works it cites.
U.S. Department of Health and Human Services: Health Information Security. Accessed: 2024-11-04 (2024). https://www.hhs.gov/hipaa/for-professionals/security/index.html
2024
Later among the works it cites.
U.S. Food and Drug Administration: 510(k) Clearances. Accessed: 2024-11-04 (2024). https://www.fda.gov/medical-devices/device-approvals-and-clearances/510k-clearances
2024
Later among the works it cites.
U.S. Food and Drug Administration: Good Machine Learning Practice for Medical Device Development: Guiding Principles. Accessed: 2024-11-04 (2024). https://www.fda.gov/medical-devices/software-medical-device-samd/good-machine-learning-practice-medical-device-development-guiding-principles
2024
Later among the works it cites.
U.S. Food and Drug Administration: Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/marketing-submission-recommendations-predetermined-change-control-plan-artificial-intelligence . Accessed: 2024-12-09 (2024)
2024
Later among the works it cites.
U.S. Food and Drug Administration: Premarket Approval (PMA). Accessed: 2024-11-04 (2024). https://www.fda.gov/medical-devices/premarket-submissions-selecting-and-preparing-correct-submission/premarket-approval-pma
2024
Later among the works it cites.
U.S. Food and Drug Administration: Software as a Medical Device (SaMD). Accessed: 2024-11-04 (2024). https://www.fda.gov/medical-devices/digital-health-center-excellence/software-medical-device-samd
2024
Later among the works it cites.
Vishwanath, P.R., Tiwari, S., Naik, T.G., Gupta, S., Thai, D.N., Zhao, W., KWON, S., Ardulov, V., Tarabishy, K., McCallum, A., et al
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Wang, D., Liang, J., Ye, J., Li, J., Li, J., Zhang, Q., Hu, Q., Pan, C., Wang, D., Liu, Z., et al
2024
Later among the works it cites.
Wu, C., Lin, W., Zhang, X., Zhang, Y., Xie, W., Wang, Y.: Pmc-llama: toward building open-source language models for medicine. Journal of the American Medical Informatics Association, 045 (2024)
2024
Later among the works it cites.
2024
Later among the works it cites.
Wang, C., Ong, J., Wang, C., Ong, H., Cheng, R., Ong, D.: Potential for gpt technology to optimize future clinical decision-making using retrieval-augmented generation. Annals of Biomedical Engineering 52
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Wang, R., Yang, Z., Zhao, Z., Tong, X., Hong, Z., Qian, K.: Llm-based robot task planning with exceptional handling for general purpose service robots. In: 2024 43rd Chinese Control Conference (CCC), pp. 4439–4444 (2024). IEEE
2024
Later among the works it cites.
Wang, D., Zhang, S.: Large language models in medical and healthcare fields: applications, advances, and challenges. Artificial Intelligence Review 57
2024
Later among the works it cites.
Xiong, G., Jin, Q., Lu, Z., Zhang, A.: Benchmarking retrieval-augmented generation for medicine. In: Ku, L.-W., Martins, A., Srikumar, V. (eds.) Findings of the Association for Computational Linguistics: ACL 2024, pp. 6233–6251. Association for Computational Linguistics, Bangkok, Thailand (2024). https://doi.org/10.18653/v1/2024.findings-acl.372 . https://aclanthology.org/2024.findings-acl.372/
2024
Later among the works it cites.
Xiong, G., Jin, Q., Wang, X., Zhang, M., Lu, Z., Zhang, A.: Improving retrieval-augmented generation in medicine with iterative follow-up questions. In: Biocomputing 2025: Proceedings of the Pacific Symposium, pp. 199–214 (2024). World Scientific
2024
Later among the works it cites.
Xia, W., Li, D., He, W., Pickhardt, P.J., Jian, J., Zhang, R., Zhang, J., Song, R., Tong, T., Yang, X., Gao, X., Cui, Y.: Multicenter evaluation of a weakly supervised deep learning model for lymph node diagnosis in rectal cancer at mri. Radiology: Artificial Intelligence (2024) https://doi.org/10.1148/ryai.230152
2024
Later among the works it cites.
2024
Later among the works it cites.
Xie, S.M., Zhang, M., Huang, M., Shi, R., Guo, L., Peng, C., Yan, P., Zhou, Y., Qiu, X.: Calibrating the confidence of large language models by eliciting fidelity. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP) (2024). https://doi.org/10.18653/v1/2024.emnlp-main.173 . Association for Computational Linguistics
2024
Later among the works it cites.
Xu, D., Zhang, Z., Zhu, Z., Lin, Z., Liu, Q., Wu, X., Xu, T., Wang, W., Ye, Y., Zhao, X., et al
2024
Later among the works it cites.
2024
Later among the works it cites.
Xu, H., Zhu, Z., Zhang, S., Ma, D., Fan, S., Chen, L., Yu, K.: Rejection improves reliability: Training LLMs to refuse unknown questions using RL from knowledge feedback. In: First Conference on Language Modeling (2024). https://openreview.net/forum?id=lJMioZBoR8
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Zheng, M., Pei, J., Logeswaran, L., Lee, M., Jurgens, D.: When ”a helpful assistant” is not really helpful: Personas in system prompts do not improve performances of large language models. In: Al-Onaizan, Y., Bansal, M., Chen, Y.-N. (eds.) Findings of the Association for Computational Linguistics: EMNLP 2024, pp. 15126–15154. Association for Computational Linguistics, Miami, Florida, USA (2024). https://doi.org/10.18653/v1/2024.findings-emnlp.888 . https://aclanthology.org/2024.findings-emnlp.888/
2024
Later among the works it cites.
2024
Later among the works it cites.
Asgari, E., Montaña-Brown, N., Dubois, M., Khalil, S., Balloch, J., Au Yeung, J., Pimenta, D.: A framework to assess clinical safety and hallucination rates of llms for medical text summarisation. npj Digital Medicine 8
2025
Closest in time.
Burke, G., Schellmann, H.: Researchers say an ai-powered transcription tool used in hospitals invents things no one ever said. AP News (2024). Accessed: 2025-02-22
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
OpenAI: GPT-5 System Card. https://cdn.openai.com/gpt-5-system-card.pdf . OpenAI model release documentation (2025)
2025
Closest in time.
2025
Closest in time.