Fetching the paper…
Reading the bibliography…
Artificial intelligence holds great promise for expanding access to expert medical knowledge and reasoning.
Experience with a model of sequential diagnosis
G Anthony Gorry and G Octo Barnett · 1968
Earlier work this paper cites.
Diagnostic strategies in the hypothesis-directed pathfinder system
Eric J Horvitz, David E Heckerman, Bharat N Nathwani, and Lawrence M Fagan · 1984
Earlier work this paper cites.
Decision theory in expert systems and artificial intelligence
Eric J Horvitz, John S Breese, and Max Henrion · 1988
Earlier work this paper cites.
Toward normative expert systems: Part I The Pathfinder project
David E Heckerman, Eric J Horvitz, and Bharat N Nathwani · 1992
Earlier work this paper cites.
Time-critical action: Representations and application
Eric Horvitz and Adam Seiver · 1997
Earlier work this paper cites.
The triple aim: care, health, and cost
Donald M Berwick, Thomas W Nolan, and John Whittington · 2008
Earlier work this paper cites.
Improving quality and curbing health care spending: opportunities for the congress and the obama administration. dartmouth atlas white paper. 2008
John E Wennberg, S Brownle, Elliot S Fisher, Jonathan S Skinner, and James N Weinstein · 2008
Earlier work this paper cites.
BioMedBERT: A pre-trained biomedical language model for QA and IR
Souradip Chakraborty, Ekaba Bisong, Shweta Bhatt, Thomas Wagner, Riley Elliott, and Francesco Mosconi · 2020
Earlier work this paper cites.
Domain-specific language model pretraining for biomedical natural language processing
Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon · 2021
Cited alongside, same era.
Comparing chatgpt and gpt-4 performance in usmle soft skill assessments
Dana Brin, Vera Sorin, Akhil Vaid, Ali Soroush, Benjamin S Glicksberg, Alexander W Charney, Girish Nadkarni, and Eyal Klang · 2023
Cited alongside, same era.
How does chatgpt perform on the united states medical licensing examination (usmle)? the implications of large language models for medical education and knowledge assessment
Aidan Gilson, Conrad W Safranek, Thomas Huang, Vimig Socrates, Ling Chi, Richard Andrew Taylor, David Chartash, et al · 2023
Cited alongside, same era.
Large language models encode clinical knowledge
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al · 2023
Cited alongside, same era.
Superhuman performance of a large language model on the reasoning tasks of a physician
From Medprompt to o1: Exploration of run-time strategies for medical challenge problems and beyond
Harsha Nori, Naoto Usuyama, Nicholas King, Scott Mayer McKinney, Xavier Fernandes, Sheng Zhang, and Eric Horvitz · 2024
Later among the works it cites.
Agentclinic: a multimodal agent benchmark to evaluate ai in simulated clinical environments
Samuel Schmidgall, Rojin Ziaei, Carl Harris, Eduardo Reis, Jeffrey Jopling, and Michael Moor · 2024
Later among the works it cites.
Healthbench: Evaluating large language models towards improved human health
Rahul K Arora, Jason Wei, Rebecca Soskin Hicks, Preston Bowman, Joaquin Quiñonero-Candela, Foivos Tsimpourlas, Michael Sharman, Meghan Shah, Andrea Vallone, Alex Beutel, et al · 2025
Closest in time.
Case 15-2025: A 52-year-old man with fever, nausea, and respiratory failure
Martín Hunter, Ignacio Lopez Saubidet, Tomás Amerio, and Maria V. Leone · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Peter G Brodeur, Thomas A Buckley, Zahir Kanjee, Ethan Goh, Evelyn Bin Ling, Priyank Jain, Stephanie Cabral, Raja-Elie Abdulnour, Adrian Haimovich, Jason A Freed, et al · 2024
Cited alongside, same era.
Clinical reasoning of a generative artificial intelligence model compared with physicians
Stephanie Cabral, Daniel Restrepo, Zahir Kanjee, Philip Wilson, Byron Crowe, Raja-Elie Abdulnour, and Adam Rodman · 2024
Cited alongside, same era.
Large language model influence on diagnostic reasoning: a randomized clinical trial
Ethan Goh, Robert Gallo, Jason Hom, Eric Strong, Yingjie Weng, Hannah Kerman, Joséphine A Cool, Zahir Kanjee, Andrew S Parsons, Neera Ahuja, et al · 2024
Cited alongside, same era.
Mediq: Question-asking llms and a benchmark for reliable interactive clinical reasoning
Stella Li, Vidhisha Balachandran, Shangbin Feng, Jonathan Ilgen, Emma Pierson, Pang Wei W Koh, and Yulia Tsvetkov · 2024
Cited alongside, same era.
Medhelm: Holistic evaluation of large language models for medical tasks
Suhana Bedi, Hejie Cui, Miguel Fuentes, Alyssa Unell, Michael Wornow, Juan M Banda, Nikesh Kotecha, Timothy Keyes, Yifan Mai, Mert Oez, et al
Cited in the paper.
The optimization paradox in clinical ai multi-agent systems
Suhana Bedi, Iddah Mlauzi, Daniel Shin, Sanmi Koyejo, and Nigam H Shah
Cited in the paper.
Capabilities of GPT-4 on medical challenge problems
Harsha Nori, Nicholas King, Scott Mayer McKinney, Dean Carignan, and Eric Horvitz
Cited in the paper.
Can generalist foundation models outcompete special-purpose tuning? Case study in medicine
Harsha Nori, Yin Tat Lee, Sheng Zhang, Dean Carignan, Richard Edgar, Nicolo Fusi, Nicholas King, Jonathan Larson, Yuanzhi Li, Weishung Liu, et al
Cited in the paper.
How ai could reshape health care—rise in direct-to-consumer models
Kenneth D Mandl · 2025
Closest in time.
Towards accurate differential diagnosis with large language models
Daniel McDuff, Mike Schaekermann, Tao Tu, Anil Palepu, Amy Wang, Jake Garrison, Karan Singhal, Yash Sharma, Shekoofeh Azizi, Kavita Kulkarni, et al · 2025
Closest in time.
Towards conversational diagnostic artificial intelligence
Tao Tu, Mike Schaekermann, Anil Palepu, Khaled Saab, Jan Freyberg, Ryutaro Tanno, Amy Wang, Brenna Li, Mohamed Amin, Yong Cheng, et al · 2025
Closest in time.