Fetching the paper…
Reading the bibliography…
We present HealthBench, an open-source benchmark measuring the performance and safety of large language models in healthcare.
Scaling-up empirical risk minimization: Optimization of incomplete u u -statistics
S. Clémençon, I. Colin, and A. Bellet · 2016
Earlier work this paper cites.
Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs
V. Gulshan, L. Peng, M. Coram, M. C. Stumpe, D. Wu, A. Narayanaswamy, S. Venugopalan, K. Widner, T. Madams, J. Cuadros, R. Kim, R. Raman, P. C. Nelson, J. L. Mega, and D. R. Webster · 2016
Earlier work this paper cites.
Big data and machine learning in health care
A. L. Beam and I. S. Kohane · 2017
Earlier work this paper cites.
Dermatologist-level classification of skin cancer with deep neural networks
A. Esteva, B. Kuprel, R. A. Novoa, J. Ko, S. M. Swetter, H. M. Blau, and S. Thrun · 2017
Earlier work this paper cites.
Pubmedqa: A dataset for biomedical research question answering
Q. Jin, D. Dhingra, and W. Cohen · 2019
Earlier work this paper cites.
High-performance medicine: The convergence of human and artificial intelligence
E. J. Topol · 2019
Earlier work this paper cites.
Radgraph: Extracting clinical entities and relations from radiology reports
S. Jain, A. Agrawal, A. Saporta, S. Q. Truong, D. N. Duong, T. Bui, P. Chambon, Y. Zhang, M. P. Lungren, A. Y. Ng, et al · 2021
Earlier work this paper cites.
What disease does this patient have? a large-scale open-domain medical qa dataset
D. Jin et al · 2021
Earlier work this paper cites.
Biogpt: generative pre-trained transformer for biomedical text generation and mining
R. Luo, L. Sun, Y. Xia, T. Qin, S. Zhang, H. Poon, and T.-Y. Liu · 2022
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, X. Jiang, K. Cobbe, T. Eloundou, G. Krueger, K. Button, M. Knight, B. Chess, and J. Schulman · 2022
Earlier work this paper cites.
Medmcqa: A large-scale multi-subject multi-choice dataset for the medical domain
S. Pal et al · 2022
Earlier work this paper cites.
Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum
J. W. Ayers, A. Poliak, M. Dredze, E. C. Leas, Z. Zhu, J. B. Kelley, D. J. Faix, A. M. Goodman, C. A. Longhurst, M. Hogarth, et al · 2023
Earlier work this paper cites.
The future landscape of large language models in medicine
J. Clusmann, F. R. Kolbinger, H. S. Muti, Z. I. Carrero, J.-N. Eckardt, et al · 2023
Earlier work this paper cites.
Evaluation of gpt-3.5 and gpt-4 for supporting real-world information needs in healthcare delivery
D. Dash, R. Thapa, J. M. Banda, A. Swaminathan, M. Cheatham, M. Kashyap, N. Kotecha, J. H. Chen, S. Gombar, L. Downing, et al · 2023
Earlier work this paper cites.
How does chatgpt perform on the united states medical licensing examination (usmle)? the implications of large language models for medical education and knowledge assessment
A. Gilson, C. W. Safranek, T. Huang, V. Socrates, L. Chi, R. A. Taylor, D. Chartash, et al · 2023
Earlier work this paper cites.
Accuracy of a generative artificial intelligence model in a complex diagnostic challenge
Z. Kanjee, B. Crowe, and A. Rodman · 2023
Earlier work this paper cites.
Benefits, limits, and risks of GPT‑4 as an AI chatbot for medicine
P. Lee, S. Bubeck, and J. Petro · 2023
Earlier work this paper cites.
Llava-med: Training a large language-and-vision assistant for biomedicine in one day
C. Li, C. Wong, S. Zhang, N. Usuyama, H. Liu, J. Yang, T. Naumann, H. Poon, and J. Gao · 2023
Cited alongside, same era.
Large language models are few‑shot health learners
X. Liu, D. McDuff, G. Kovacs, I. Galatzer‑Levy, J. Sunshine, et al · 2023
Cited alongside, same era.
Med‑flamingo: A multimodal medical few‑shot learner
M. Moor, L. von Rueden, S. Adler, W. Ping, H. Valentin, et al · 2023
Cited alongside, same era.
Capabilities of gpt-4 on medical challenge problems
H. Nori, N. King, S. McKinney, D. Carignan, and E. Horvitz · 2023
Cited alongside, same era.
Large language models encode clinical knowledge
K. Singhal, S. Azizi, T. Tu, S. S. Mahdavi, J. Wei, H. W. Chung, N. Scales, A. Tanwani, H. Cole-Lewis, S. Pfohl, et al · 2023
Cited alongside, same era.
Capabilities of gemini models in medicine
K. Saab, T. Tu, W.-H. Weng, R. Tanno, D. Stutz, E. Wulczyn, F. Zhang, T. Strother, C. Park, E. Vedadi, et al · 2024
Later among the works it cites.
Towards generalist biomedical ai
T. Tu, S. Azizi, D. Driess, M. Schaekermann, M. Amin, P.-C. Chang, A. Carroll, C. Lau, R. Tanno, I. Ktena, et al · 2024
Later among the works it cites.
Physician‑ and large language model‑generated hospital discharge summaries: A blinded, comparative quality and safety study
C. Y. K. Williams, C. R. Subramanian, S. S. Ali, and et al · 2024
Later among the works it cites.
Justice or prejudice? quantifying biases in llm-as-a-judge
J. Ye, Y. Wang, Y. Huang, D. Chen, Q. Zhang, N. Moniz, T. Gao, W. Geyer, C. Huang, P.-Y. Chen, et al · 2024
Later among the works it cites.
Almanac — retrieval‑augmented language models for clinical medicine
C. Zakka, N. Kiani, A. Wong, J. Shi, C. Raffel, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Performance of gpt-3.5 and gpt-4 on the japanese medical licensing examination: comparison study
S. Takagi, T. Watari, A. Erabi, K. Sakaguchi, et al · 2023
Cited alongside, same era.
Large language models in medicine
A. J. Thirunavukarasu, D. S. J. Ting, K. Elangovan, L. Gutierrez, T. F. Tan, and D. S. W. Ting · 2023
Cited alongside, same era.
Chatgpt’s ability to assist with clinical documentation: A randomized controlled trial
H. P. Baker, E. Dwyer, S. Kalidoss, and et al · 2024
Cited alongside, same era.
Humans or llms as the judge? a study on judgement biases
G. H. Chen, S. Chen, Z. Liu, F. Jiang, and B. Wang · 2024
Cited alongside, same era.
Towards a personal health large language model
J. Cosentino, A. Belyaeva, X. Liu, N. A. Furlotte, Z. Yang, C. Lee, E. Schenck, Y. Patel, J. Cui, L. D. Schneider, et al · 2024
Cited alongside, same era.
Autonomous medical evaluation for guideline adherence of large language models
D. Fast, L. C. Adams, F. Busch, C. Fallon, M. Huppertz, R. Siepmann, P. Prucker, N. Bayerl, D. Truhn, M. Makowski, et al · 2024
Cited alongside, same era.
Medalign: A clinician-generated dataset for instruction following with electronic medical records
S. L. Fleming, A. Lozano, W. J. Haberkorn, J. A. Jindal, E. Reis, R. Thapa, L. Blankemeier, J. Z. Genkins, E. Steinberg, A. Nayak, et al · 2024
Cited alongside, same era.
Later among the works it cites.
Wildchat: 1m chatgpt interaction logs in the wild
W. Zhao, X. Ren, J. Hessel, C. Cardie, Y. Choi, and Y. Deng · 2024
Later among the works it cites.
Lmsys-chat-1m: A large-scale real-world llm conversation dataset
L. Zheng, W.-L. Chiang, Y. Sheng, T. Li, S. Zhuang, Z. Wu, Y. Zhuang, Z. Li, Z. Lin, E. P. Xing, J. E. Gonzalez, I. Stoica, and H. Zhang · 2024
Later among the works it cites.
Randomized trial of a generative ai chatbot for mental health treatment
M. V. Heinz, N. Jacobson, E. Wright, and et al · 2025
Closest in time.
Wildbench: Benchmarking LLMs with challenging tasks from real users in the wild
B. Y. Lin, Y. Deng, K. Chandu, A. Ravichander, V. Pyatkin, N. Dziri, R. Le Bras, and Y. Choi · 2025
Closest in time.
Towards accurate differential diagnosis with large language models
D. McDuff, M. Schaekermann, T. Tu, and et al · 2025
Closest in time.
VISTA: Visual–language understanding leaderboard
Scale AI · 2025
Closest in time.
Holistic evaluation of large language models for medical applications, 2025
N. Shah, M. Pfeffer, and P. Liang · 2025
Closest in time.
Toward expert-level medical question answering with large language models
K. Singhal, T. Tu, J. Gottweis, R. Sayres, E. Wulczyn, M. Amin, L. Hou, K. Clark, S. R. Pfohl, H. Cole-Lewis, et al · 2025
Closest in time.
V. Sirdeshmukh, K. Deshpande, J. Mols, L. Jin, E.-Y. Cardona, D. Lee, J. Kritz, W. Primack, S. Yue, and C. Xing · 2025
Closest in time.
Paperbench: Evaluating ai’s ability to replicate ai research
G. Starace, O. Jaffe, D. Sherburn, J. Aung, J. S. Chan, L. Maksin, R. Dias, E. Mays, B. Kinsella, W. Thompson, et al · 2025
Closest in time.
Collaboration between clinicians and vision–language models in radiology report generation
R. Tanno, D. G. Barrett, A. Sellergren, S. Ghaisas, S. Dathathri, A. See, J. Welbl, C. Lau, T. Tu, S. Azizi, et al · 2025
Closest in time.
Towards conversational diagnostic artificial intelligence
T. Tu, M. Schaekermann, A. Palepu, K. Saab, J. Freyberg, R. Tanno, A. Wang, B. Li, M. Amin, Y. Cheng, et al · 2025
Closest in time.