Fetching the paper…
Reading the bibliography…
While large language models (LLMs) achieve near-perfect scores on medical licensing exams, these evaluations inadequately reflect the complexity and diversity of real-world clinical practice.
2009
Earlier work this paper cites.
Guha, B.: Secret ballots and costly information gathering: the jury size problem revisited. MPRA Paper No. 73048 (2016). https://mpra.ub.uni-muenchen.de/73048/
2016
Earlier work this paper cites.
Pal, A., Umapathi, L.K., Sankarasubbu, M.: Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering 174
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., Newman, B., Yuan, B., Yan, B., Zhang, C., Cosgrove, C., Manning, C.D., Ré, C., Acosta-Navas, D., Hudson, D.A., Zelikman, E., Durmus, E., Ladhak, F., Rong, F., Ren, H., Yao, H., Wang, J., Santhanam, K., Orr, L., Zheng, L., Yuksekgonul, M., Suzgun, M., Kim, N., Guha, N., Chatterji, N., Khattab, O., Henderson, P., Huang, Q., Chi, R., Xie, S.M., Santurkar, S., Ganguli, S., Hashimoto, T., Icard, T., Zhang, T., Chaudhary, V., Wang, W., Li, X., Mai, Y., Zhang, Y., Koreeda, Y.: Holistic evaluation of language models. Transactions on Machine Learning Research TMLR
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Papers with Code: Question Answering on MedQA (USMLE) - Papers with Code. https://paperswithcode.com/sota/question-answering-on-medqa-usmle . Accessed: 2025-04-22 (2024)
2024
Cited alongside, same era.
Khosravi, M., Zare, Z., Mojtabaeian, S.M., Izadi, R.: Artificial intelligence and decision-making in healthcare: A thematic analysis of a systematic review of reviews. Health Services Research and Managerial Epidemiology 11
2024
Cited alongside, same era.
Nath, D.: Artificial intelligence (ai) will transform the clinical workflow with the next-generation technology. HealthTech Magazines (2024). AVP & Deputy CIO, Downstate Health Sciences University
2024
Cited alongside, same era.
Bedi, S., Liu, Y., Orr-Ewing, L., Dash, D., Koyejo, S., Callahan, A., Fries, J.A., Wornow, M., Swaminathan, A., Lehmann, L.S., Hong, H.J., Kashyap, M., Chaurasia, A.R., Shah, N.R., Singh, K., Tazbaz, T., Milstein, A., Pfeffer, M.A., Shah, N.H.: Testing and evaluation of health care applications of large language models: A systematic review. JAMA 333
Confident AI: DeepEval: Open-Source Evaluation Framework for LLMs. https://github.com/confident-ai/deepeval . Accessed: 2025-05-02 (2024). https://github.com/confident-ai/deepeval
2024
Later among the works it cites.
Van Veen, D., Van Uden, C., Blankemeier, L., Delbrouck, J.-B., Aali, A., Bluethgen, C., Pareek, A., Polacin, M., Reis, E.P., Seehofnerová, A., Rohatgi, N., Hosamani, P., Collins, W., Ahuja, N., Langlotz, C.P., Hom, J., Gatidis, S., Pauly, J., Chaudhari, A.S.: Adapted large language models can outperform medical experts in clinical text summarization. Nature Medicine 30
2024
Later among the works it cites.
Carl, N., Haggenmüller, S., Wies, C., Nguyen, L., Winterstein, J.T., Hetz, M.J., Mangold, M.H., Hartung, F.O., Grüne, B., et al
2025
Closest in time.
Raji, I.D., Daneshjou, R., Alsentzer, E.: It’s time to bench the medical exam benchmark. NEJM AI 2
2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Earlier work this paper cites.
2024
Cited alongside, same era.
Hager, P., Jungmann, F., Holland, R., Bhagat, K., Hubrecht, I., Knauer, M., Vielhauer, J., Makowski, M., Braren, R., Kaissis, G., Rueckert, D.: Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nature Medicine 30
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Qiu, P., Wu, C., Zhang, X., Lin, W., Wang, H., Zhang, Y., Wang, Y., Xie, W.: Towards building multilingual language model for medicine. Nature Communications 15
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Arora, R.K., Wei, J., Soskin Hicks, R., Bowman, P., Quiñonero Candela, J., Tsimpourlas, F., Sharman, M., Shah, M., Vallone, A., Beutel, A., Heidecke, J., Singhal, K.: HealthBench: Evaluating Large Language Models Towards Improved Human Health. https://cdn.openai.com/pdf/bd7a39d5-9e9f-47b3-903c-8b847ca650c7/healthbench_paper.pdf . Accessed 13 May 2025 (2025). https://openai.com/index/healthbench
2025
Closest in time.
Croxford, E., Gao, Y., First, E., Pellegrino, N., Schnier, M., Caskey, J., Oguss, M., Wills, G., Chen, G., Dligach, D., Churpek, M.M., Mayampurath, A., Liao, F., Goswami, C., Wong, K.K., Patterson, B.W., Afshar, M.: Automating evaluation of ai text generation in healthcare with a large language model (llm)-as-a-judge. medRxiv (2025) https://doi.org/10.1101/2025.04.22.25326219 . Preprint, not peer-reviewed
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Samuylova, E.: LLM-as-a-judge: a Complete Guide to Using LLMs for Evaluations. Evidently AI. Accessed 2025-05-02. https://www.evidentlyai.com/llm-guide/llm-as-a-judge
2025
Closest in time.