Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are currently being used to answer medical questions across a variety of clinical domains.
PubMedQA: A dataset for biomedical research question answering
Jin, Q., Dhingra, B., Liu, Z., Cohen, W. W., and Lu, X · 2019
Earlier work this paper cites.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Jin, D., Pan, E., Oufattole, N., Weng, W.-H., Fang, H., and Szolovits, P · 2020
Earlier work this paper cites.
WebGPT: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., Jiang, X., Cobbe, K., Eloundou, T., Krueger, G., Button, K., Knight, M., Chess, B., and Schulman, J · 2021
Earlier work this paper cites.
Measuring attribution in natural language generation models
Rashkin, H., Nikolaev, V., Lamm, M., Aroyo, L., Collins, M., Das, D., Petrov, S., Tomar, G. S., Turc, I., and Reitter, D · 2021
Earlier work this paper cites.
Attributed question answering: Evaluation and modeling for attributed large language models
Bohnet, B., Tran, V. Q., Verga, P., Aharoni, R., Andor, D., Soares, L. B., Ciaramita, M., Eisenstein, J., Ganchev, K., Herzig, J., Hui, K., Kwiatkowski, T., Ma, J., Ni, J., Saralegui, L. S., Schuster, T., Cohen, W. W., Collins, M., Das, D., Metzler, D., Petrov, S., and Webster, K · 2022
Earlier work this paper cites.
Teaching language models to support answers with verified quotes
Menick, J., Trebacz, M., Mikulik, V., Aslanides, J., Song, F., Chadwick, M., Glaese, M., Young, S., Campbell-Gillingham, L., Irving, G., and McAleese, N · 2022
Earlier work this paper cites.
Exploring perceptions and experiences of ChatGPT in medical education: A qualitative study among medical college faculty and students in saudi arabia
Abouammoh, N., Alhasan, K., Raina, R., Malki, K. A., Aljamaan, F., Tamimi, I., Muaygil, R., Wahabi, H., Jamal, A., Al-Tawfiq, J. A., Al-Eyadhy, A., Soliman, M., and Temsah, M.-H · 2023
Earlier work this paper cites.
Creating trustworthy LLMs: Dealing with hallucinations in healthcare AI
Ahmad, M. A., Yaramis, I., and Roy, T. D · 2023
Earlier work this paper cites.
Comparing ChatGPT and GPT-4 performance in USMLE soft skill assessments
Brin, D., Sorin, V., Vaid, A., Soroush, A., Glicksberg, B. S., Charney, A. W., Nadkarni, G., and Klang, E · 2023
Earlier work this paper cites.
Understanding retrieval augmentation for Long-Form question answering
Chen, H.-T., Xu, F., Arora, S., and Choi, E · 2023
Earlier work this paper cites.
Evaluation of GPT-3.5 and GPT-4 for supporting real-world information needs in healthcare delivery
Dash, D., Thapa, R., Banda, J. M., Swaminathan, A., Cheatham, M., Kashyap, M., Kotecha, N., Chen, J. H., Gombar, S., Downing, L., Pedreira, R., Goh, E., Arnaout, A., Morris, G. K., Magon, H., Lungren, M. P., Horvitz, E., and Shah, N. H · 2023
Earlier work this paper cites.
Use of GPT-4 to diagnose complex clinical cases
Eriksen Alexander V., Möller Sören, and Ryg Jesper · 2023
Earlier work this paper cites.
Enabling large language models to generate text with citations
Gao, T., Yen, H., Yu, J., and Chen, D · 2023
Earlier work this paper cites.
Large language model AI chatbots require approval as medical devices
Gilbert, S., Harvey, H., Melvin, T., Vollebregt, E., and Wicks, P · 2023
Earlier work this paper cites.
How to safely integrate large language models into health care
Gottlieb, S. and Silvis, L · 2023
Cited alongside, same era.
Regulating ChatGPT and other large generative AI models
Hacker, P., Engel, A., and Mauer, M · 2023
Cited alongside, same era.
AI-Generated medical Advice-GPT and beyond
Haupt, C. E. and Marks, M · 2023
Cited alongside, same era.
HAGRID: A Human-LLM collaborative dataset for generative Information-Seeking with attribution
Kamalloo, E., Jafari, A., Zhang, X., Thakur, N., and Lin, J · 2023
Cited alongside, same era.
Towards verifiable generation: A benchmark for knowledge-aware language model attribution
Li, X., Cao, Y., Pan, L., Ma, Y., and Sun, A · 2023
Cited alongside, same era.
Evaluating verifiability in generative search engines
Liu, N. F., Zhang, T., and Liang, P · 2023
Cited alongside, same era.
Automatic evaluation of attribution by large language models
Yue, X., Wang, B., Chen, Z., Zhang, K., Su, Y., and Sun, H · 2023
Later among the works it cites.
ChatGPT and clinical training: Perception, concerns, and practice of Pharm-D students
Zawiah, M., Al-Ashwal, F. Y., Gharaibeh, L., Abu Farha, R., Alzoubi, K. H., Abu Hammour, K., Qasim, Q. A., and Abrah, F · 2023
Later among the works it cites.
ChatGPT hallucinates when attributing answers
Zuccon, G., Koopman, B., and Shaik, R · 2023
Later among the works it cites.
ChatGPT poses new regulatory questions for FDA, medical industry
Baumann, J · 2024
Closest in time.
Large legal fictions: Profiling legal hallucinations in large language models
Dahl, M., Magesh, V., Suzgun, M., and Ho, D. E · 2024
Closest in time.
Medical chatbot using OpenAI’s GPT-3 told a fake patient to kill themselves
Daws, R · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ExpertQA: Expert-Curated questions and attributed answers
Malaviya, C., Lee, S., Chen, S., Sieber, E., Yatskar, M., and Roth, D · 2023
Cited alongside, same era.
FActScore: Fine-grained atomic evaluation of factual precision in long form text generation
Min, S., Krishna, K., Lyu, X., Lewis, M., Yih, W.-T., Koh, P. W., Iyyer, M., Zettlemoyer, L., and Hajishirzi, H · 2023
Cited alongside, same era.
Does ChatGPT provide appropriate and equitable medical advice?: A vignette-based, clinical evaluation across care contexts
Nastasi, A. J., Courtright, K. R., Halpern, S. D., and Weissman, G. E · 2023
Cited alongside, same era.
Med-HALT: Medical domain hallucination test for large language models
Pal, A., Umapathi, L. K., and Sankarasubbu, M · 2023
Cited alongside, same era.
WebCPM: Interactive web search for Chinese long-form question answering
Qin, Y., Cai, Z., Jin, D., Yan, L., Liang, S., Zhu, K., Lin, Y., Han, X., Ding, N., Wang, H., Xie, R., Qi, F., Liu, Z., Sun, M., and Zhou, J · 2023
Cited alongside, same era.
Chatbot vs medical student performance on Free-Response clinical reasoning examinations
Strong, E., DiGiammarino, A., Weng, Y., Kumar, A., Hosamani, P., Hom, J., and Chen, J. H · 2023
Cited alongside, same era.
Closest in time.
A boy saw 17 doctors over 3 years for chronic pain. ChatGPT found the diagnosis
Holohan, M · 2024
Closest in time.
ChatGPT used by mental health tech app in AI experiment with users
Ingram, D · 2024
Closest in time.
Large language models in medicine: The potential to reduce workloads, leverage the EMR for better communication & more
Jansz, J. and Sadelski, P. T · 2024
Closest in time.
Loneliness and suicide mitigation for students using GPT3-enabled chatbots
Maples, B., Cerit, M., Vishwanath, A., and Pea, R · 2024
Closest in time.
TrustLLM: Trustworthiness in large language models
Sun, L., Huang, Y., Wang, H., Wu, S., Zhang, Q., Gao, C., Huang, Y., Lyu, W., Zhang, Y., Li, X., Liu, Z., Liu, Y., Wang, Y., Zhang, Z., Kailkhura, B., Xiong, C., Xiao, C., Li, C., Xing, E., Huang, F., Liu, H., Ji, H., Wang, H., Zhang, H., Yao, H., Kellis, M., Zitnik, M., Jiang, M., Bansal, M., Zou, J., Pei, J., Liu, J., Gao, J., Han, J., Zhao, J., Tang, J., Wang, J., Mitchell, J., Shu, K., Xu, K., Chang, K.-W., He, L., Huang, L., Backes, M., Gong, N. Z., Yu, P. S., Chen, P.-Y., Gu, Q., Xu, R., Ying, R., Ji, S., Jana, S., Chen, T., Liu, T., Zhou, T., Wang, W., Li, X., Zhang, X., Wang, X., Xie, X., Chen, X., Wang, X., Liu, Y., Ye, Y., Cao, Y., Chen, Y., and Zhao, Y · 2024
Closest in time.
AlpacaEval leaderboard
Tatsu · 2024
Closest in time.
FDA calls for ‘nimble’ regulation of ChatGPT-like models to avoid being ‘swept up quickly’ by tech
Taylor, N. P · 2024
Closest in time.