Fetching the paper…
Reading the bibliography…
While large language models (LLMs) have shown promise in diagnostic dialogue, their capabilities for effective management reasoning - including disease progression, therapeutic response, and safe medication prescription - remain under-explored.
“Human and computer-aided diagnosis of abdominal pain: further report with emphasis on performance of clinicians”
F De, D Leaper, J Horrocks, J Staniland and A McCann · 1974
Earlier work this paper cites.
“Assessment of clinical competence using objective structured examination”
R Harden, M Stevenson, W Downie and G Wilson · 1975
Earlier work this paper cites.
“Computer aided diagnosis of acute abdominal pain: a multicentre study”
I Adams, M Chan, P Clifford, W Cooke, V Dallos, F de Dombal, M Edwards, D Hancock, D Hewett and N McIntyre · 1986
Earlier work this paper cites.
“Computer programs to support clinical decision making”
Edward Shortliffe · 1987
Earlier work this paper cites.
“Cognitive processes in clinical inference and decision making.”, 1988
Arthur. Elstein · 1988
Earlier work this paper cites.
“Semantic structures and diagnostic thinking of experts and novices”
G Bordage and M Lemieux · 1991
Earlier work this paper cites.
“Clinical guidelines: potential benefits, limitations, and harms of clinical guidelines”
S Woolf, R Grol, A Hutchinson, M Eccles and J Grimshaw · 1999
Earlier work this paper cites.
“Putting them through their PACES. Practical Assessment of Clinical Examination Skills”
Helen Gentles · 2003
Earlier work this paper cites.
“Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”, 2021
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel and Douwe Kiela · 2005
Earlier work this paper cites.
“A universal model of diagnostic reasoning”
Pat Croskerry · 2009
Earlier work this paper cites.
“Effects of evidence-based clinical practice guidelines on quality of care: a systematic review”
M Lugtenberg, J Burgers and G Westert · 2009
Earlier work this paper cites.
“Thinking, fast and slow”
Daniel Kahneman · 2011
Earlier work this paper cites.
“Inter-rater reliability: Comparison of checklist and global scoring for OSCEs”
Bunmi Malau-Aduli, Sue Mulcahy, Emma Warnecke, Petr Otahal, Peta-Ann Teague, Richard Turner and Cees Vleuten · 2012
Earlier work this paper cites.
“Improving diagnosis in health care”
John Ball, Bryan Miller and Erin Balogh · 2015
Earlier work this paper cites.
“When guidelines don’t guide: the effect of patient context on management decisions based on clinical practice guidelines”
Mathew Mercuri, Jonathan Sherbino, Robert Sedran, Jason Frank, Amiram Gafni and Geoffrey Norman · 2015
Earlier work this paper cites.
“Speech recognition for medical conversations”
Chung-Cheng Chiu, Anshuman Tripathi, Katherine Chou, Chris Co, Navdeep Jaitly, Diana Jaunzeikare, Anjuli Kannan, Patrick Nguyen, Hasim Sak and Ananth Sankar · 2017
Earlier work this paper cites.
“Does an objective structured clinical examination fit your assessment toolbox?”
Dotun Ogunyemi and Denise Dupras · 2017
Earlier work this paper cites.
“Management reasoning: Beyond the diagnosis”
David Cook, Jonathan Sherbino and Steven Durning · 2018
Earlier work this paper cites.
“It’s the destination: diagnostic accuracy and reasoning”
Sandra Monteiro, Jonathan Sherbino, Henk Schmidt, Silvia Mamede, Jonathan Ilgen and Geoff Norman · 2020
Earlier work this paper cites.
“The TeleHealth OSCE: Preparing trainees to use telemedicine as a tool for transitions of care”
Daniel Sartori, Rachael Hayes, Margaret Horlick, Jennifer Adams and Sondra Zabar · 2020
Earlier work this paper cites.
“What disease does this patient have? a large-scale open domain question answering dataset from medical exams”
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang and Peter Szolovits · 2021
Earlier work this paper cites.
“Mortality and guideline-directed medical therapy in real-world heart failure patients with reduced ejection fraction”
Peter McCullough, Hirsch Mehta, Colin Barker, Joanna Van, Sarah Mollenkopf, Candace Gunnarsson, Michael Ryan and David Cork · 2021
Earlier work this paper cites.
“Drug shortage: Causes, impact, and mitigation strategies”
Sundus Shukar, Fatima Zahoor, Khezar Hayat, Amna Saeed, Ali Gillani, Sumaira Omer, Shuchen Hu, Zaheer-Ud-Din Babar, Yu Fang and Caijun Yang · 2021
Earlier work this paper cites.
“The impact of digital patient portals on health outcomes, system efficiency, and patient attitudes: Updated systematic literature review”
Elettra Carini, Leonardo Villani, Angelo Pezzullo, Andrea Gentili, Andrea Barbara, Walter Ricciardi and Stefania Boccia · 2021
Earlier work this paper cites.
“What disease does this patient have? a large-scale open domain question answering dataset from medical exams”
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang and Peter Szolovits · 2021
Earlier work this paper cites.
“Training language models to follow instructions with human feedback”
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama and Alex Ray · 2022
Earlier work this paper cites.
“Self-consistency improves chain of thought reasoning in language models”
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi and Denny Zhou · 2022
Earlier work this paper cites.
“Do As I Can, Not As I Say: Grounding Language in Robotic Affordances”, 2022
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Ruano, Kyle Jeffrey, Sally Jesmonth, Nikhil Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Kuang-Huei Lee, Sergey Levine, Yao Lu, Linda Luu, Carolina Parada, Peter Pastor, Jornell Quiambao, Kanishka Rao, Jarek Rettinghouse, Diego Reyes, Pierre Sermanet, Nicolas Sievers, Clayton Tan, Alexander Toshev, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Mengyuan Yan and Andy Zeng · 2022
Earlier work this paper cites.
“Chain-of-thought prompting elicits reasoning in large language models”
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc Le and Denny Zhou · 2022
Earlier work this paper cites.
“Large Language Models are Zero-Shot Reasoners”
Takeshi Kojima, Shixiang Gu, Machel Reid, Yutaka Matsuo and Yusuke Iwasawa · 2022
Earlier work this paper cites.
“The state of telehealth before and after the COVID-19 pandemic”
Julia Shaver · 2022
Earlier work this paper cites.
“Chain-of-thought prompting elicits reasoning in large language models”
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc Le and Denny Zhou · 2022
Earlier work this paper cites.
“Large Language Models are Zero-Shot Reasoners”
Takeshi Kojima, Shixiang Gu, Machel Reid, Yutaka Matsuo and Yusuke Iwasawa · 2022
Earlier work this paper cites.
“Training language models to follow instructions with human feedback”
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama and Alex Ray · 2022
Earlier work this paper cites.
“Accuracy of a Generative Artificial Intelligence Model in a Complex Diagnostic Challenge”
Zahir Kanjee, Byron Crowe and Adam Rodman · 2023
Earlier work this paper cites.
“Towards Accurate Differential Diagnosis with Large Language Models”
Daniel McDuff, Mike Schaekermann, Tao Tu, Anil Palepu, Amy Wang, Jake Garrison, Karan Singhal, Yash Sharma, Shekoofeh Azizi and Kavita Kulkarni · 2023
Earlier work this paper cites.
“Gemini: a family of highly capable multimodal models”
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew Dai and Anja Hauth · 2023
Cited alongside, same era.
Rohan Anil, Andrew Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey and Zhifeng Chen · 2023
Cited alongside, same era.
“Can large language models reason about medical questions?”, 2023
Valentin Liévin, Christoffer Hother, Andreas Motzfeldt and Ole Winther · 2023
Cited alongside, same era.
“Towards expert-level medical question answering with large language models”
Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Le Hou, Kevin Clark, Stephen Pfohl, Heather Cole-Lewis and Darlene Neal · 2023
Cited alongside, same era.
“Use of a large language model to assess clinical acuity of adults in the emergency department”
Christopher Williams, Travis Zack, Brenda Miao, Madhumita Sushil, Michelle Wang, Aaron Kornblith and Atul Butte · 2024
Later among the works it cites.
“Assessing large language models for oncology data inference from radiology reports”
Li-Ching Chen, Travis Zack, Arda Demirci, Madhumita Sushil, Brenda Miao, Corynn Kasap, Atul Butte, Eric Collisson and Julian Hong · 2024
Later among the works it cites.
“Evaluating the use of large language models to provide clinical recommendations in the Emergency Department”
Christopher Williams, Brenda Miao, Aaron Kornblith and Atul Butte · 2024
Later among the works it cites.
“MedCalc-Bench: Evaluating large language models for medical calculations”
Nikhil Khandekar, Qiao Jin, Guangzhi Xiong, Soren Dunn, Serina Applebaum, Zain Anwar, Maame Sarfo-Gyamfi, Conrad Safranek, Abid Anwar, Andrew Zhang, Aidan Gilson, Maxwell Singer, Amisha Dave, Andrew Taylor, Aidong Zhang, Qingyu Chen and Zhiyong Lu · 2024
Later among the works it cites.
“Almanac - retrieval-augmented language models for clinical medicine”
Cyril Zakka, Rohan Shad, Akash Chaurasia, Alex Dalal, Jennifer Kim, Michael Moor, Robyn Fong, Curran Phillips, Kevin Alexander, Euan Ashley, Jack Boyd, Kathleen Boyd, Karen Hirsch, Curt Langlotz, Rita Lee, Joanna Melia, Joanna Nelson, Karim Sallam, Stacey Tullis, Melissa Vogelsong, John Cunningham and William Hiesinger · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Management reasoning: empirical determination of key features and a conceptual model”
David Cook, Christopher Stephenson, Larry Gruppen and Steven Durning · 2023
Cited alongside, same era.
“Diagnostic accuracy of artificial intelligence in virtual primary care”
Dan Zeltzer, Lee Herzog, Yishai Pickman, Yael Steuerman, Ran Ber, Zehavi Kugler, Ran Shaul and Jon Ebbert · 2023
Cited alongside, same era.
“Generative Agents: Interactive Simulacra of Human Behavior”, 2023
Joon Park, Joseph. O’Brien, Carrie. Cai, Meredith Morris, Percy Liang and Michael. Bernstein · 2023
Cited alongside, same era.
Alexander Vezhnevets, John. Agapiou, Avia Aharon, Ron Ziv, Jayd Matyas, Edgar. Duéñez-Guzmán, William. Cunningham, Simon Osindero, Danny Karmon and Joel. Leibo · 2023
Cited alongside, same era.
“HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face”, 2023
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu and Yueting Zhuang · 2023
Cited alongside, same era.
“Voyager: An Open-Ended Embodied Agent with Large Language Models”, 2023
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan and Anima Anandkumar · 2023
Cited alongside, same era.
“Self-refine: Iterative refinement with self-feedback”
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye and Yiming Yang · 2023
Cited alongside, same era.
“Large language models encode clinical knowledge”
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Mahdavi, Jason Wei, Hyung Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis and Stephen Pfohl · 2023
Cited alongside, same era.
Later among the works it cites.
“Large language models in medicine: A review of current clinical trials across healthcare applications”
Mahmud Omar, Girish Nadkarni, Eyal Klang and Benjamin Glicksberg · 2024
Later among the works it cites.
“Polaris: A safety-focused LLM constellation architecture for healthcare”
Subhabrata Mukherjee, Paul Gamble, Markel Ausin, Neel Kant, Kriti Aggarwal, Neha Manjunath, Debajyoti Datta, Zhengliang Liu, Jiayuan Ding, Sophia Busacca, Cezanne Bianco, Swapnil Sharma, Rae Lasko, Michelle Voisard, Sanchay Harneja, Darya Filippova, Gerry Meixiong, Kevin Cha, Amir Youssefi, Meyhaa Buvanesh, Howard Weingram, Sebastian Bierman-Lytle, Harpreet Mangat, Kim Parikh, Saad Godil and Alex Miller · 2024
Later among the works it cites.
“Conversational medical AI: Ready for practice”
Antoine Lizée, Pierre-Auguste Beaucoté, James Whitbeck, Marion Doumeingts, Anaël Beaugnon and Isabelle Feldhaus · 2024
Later among the works it cites.
“Polaris: A Safety-focused LLM Constellation Architecture for Healthcare”
Subhabrata Mukherjee, Paul Gamble, Markel Ausin, Neel Kant, Kriti Aggarwal, Neha Manjunath, Debajyoti Datta, Zhengliang Liu, Jiayuan Ding and Sophia Busacca · 2024
Later among the works it cites.
“Agents thinking fast and slow: A talker-reasoner architecture”
Konstantina Christakopoulou, Shibl Mourad and Maja Matarić · 2024
Later among the works it cites.
Zihao Wang, Shaofei Cai, Guanzhou Chen, Anji Liu, Xiaojian Ma and Yitao Liang · 2024
Later among the works it cites.
“Self-Discover: Large Language Models Self-Compose Reasoning Structures”, 2024
Pei Zhou, Jay Pujara, Xiang Ren, Xinyun Chen, Heng-Tze Cheng, Quoc. Le, Ed. Chi, Denny Zhou, Swaroop Mishra and Huaixiu Zheng · 2024
Later among the works it cites.
“Gemini 2.0 Flash Thinking” accessed: 2025-02-28, 2024
Google · 2024
Later among the works it cites.
“Learning to reason with LLMs” accessed: 2025-02-28, 2024
OpenAI · 2024
Later among the works it cites.
“Do patients who read visit notes on the patient portal have a higher rate of “loop closure” on diagnostic tests and referrals in primary care? A retrospective cohort study”
Sigall Bell, Maelys Amat, Timothy Anderson, Mark Aronson, James Benneyan, Leonor Fernandez, Dru Ricci, Talya Salant, Gordon Schiff, Umber Shafiq, Sara Singer, Scot Sternberg, Cancan Zhang and Russell Phillips · 2024
Later among the works it cites.
“Superhuman performance of a large language model on the reasoning tasks of a physician”
Peter Brodeur, Thomas Buckley, Zahir Kanjee, Ethan Goh, Evelyn Ling, Priyank Jain, Stephanie Cabral, Raja-Elie Abdulnour, Adrian Haimovich, Jason Freed, Andrew Olson, Daniel Morgan, Jason Hom, Robert Gallo, Eric Horvitz, Jonathan Chen, Arjun Manrai and Adam Rodman · 2024
Later among the works it cites.
“Towards conversational diagnostic ai”
Tao Tu, Anil Palepu, Mike Schaekermann, Khaled Saab, Jan Freyberg, Ryutaro Tanno, Amy Wang, Brenna Li, Mohamed Amin and Nenad Tomasev · 2024
Later among the works it cites.
“Automata-based constraints for language model decoding”
Terry Koo, Frederick Liu and Luheng He · 2024
Later among the works it cites.
“"We Need Structured Output": Towards User-centered Constraints on Large Language Model Output”
Michael Liu, Frederick Liu, Alexander Fiannaca, Terry Koo, Lucas Dixon, Michael Terry and Carrie Cai · 2024
Later among the works it cites.
“Towards Democratization of Subspeciality Medical Expertise”
Jack O’Sullivan, Anil Palepu, Khaled Saab, Wei-Hung Weng, Yong Cheng, Emily Chu, Yaanik Desai, Aly Elezaby, Daniel Kim and Roy Lan · 2024
Later among the works it cites.
“Exploring Large Language Models for Specialist-level Oncology Care”
Anil Palepu, Vikram Dhillon, Polly Niravath, Wei-Hung Weng, Preethi Prasad, Khaled Saab, Ryutaro Tanno, Yong Cheng, Hanh Mai and Ethan Burns · 2024
Later among the works it cites.
“Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context”, 2024
Gemini Team et al · 2024
Later among the works it cites.
“The Top 300 of 2022, ClinCalc DrugStats Database Version 2024.08”, 2024
Kane SP · 2024
Later among the works it cites.
“On context specificity and management reasoning: moving beyond diagnosis”
James Boyle, Matthew Walters, Fiona Burton, Catherine Paton, Martin Hughes, Susan Jamieson and Steven Durning · 2025
Closest in time.
“GPT-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trial”
Ethan Goh, Robert Gallo, Eric Strong, Yingjie Weng, Hannah Kerman, Jason Freed, Joséphine Cool, Zahir Kanjee, Kathleen Lane, Andrew Parsons, Neera Ahuja, Eric Horvitz, Daniel Yang, Arnold Milstein, Andrew Olson, Jason Hom, Jonathan Chen and Adam Rodman · 2025
Closest in time.
“s1: Simple test-time scaling”
Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès and Tatsunori Hashimoto · 2025
Closest in time.
“Towards an AI co-scientist”, 2025
Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, Khaled Saab, Dan Popovici, Jacob Blum, Fan Zhang, Katherine Chou, Avinatan Hassidim, Burak Gokturk, Amin Vahdat, Pushmeet Kohli, Yossi Matias, Andrew Carroll, Kavita Kulkarni, Nenad Tomasev, Yuan Guan, Vikram Dhillon, Eeshit Vaishnav, Byron Lee, Tiago Costa, José Penadés, Gary Peltz, Yunhan Xu, Annalisa Pawlosky, Alan Karthikesalingam and Vivek Natarajan · 2025
Closest in time.
“An evaluation framework for clinical use of large language models in patient interaction tasks”
Shreya Johri, Jaehwan Jeong, Benjamin Tran, Daniel Schlessinger, Shannon Wongvibulsin, Leandra Barnes, Hong-Yu Zhou, Zhuo Cai, Eliezer Van, David Kim, Roxana Daneshjou and Pranav Rajpurkar · 2025
Closest in time.
“Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents”, 2025
Junkai Li, Yunghwei Lai, Weitao Li, Jingyi Ren, Meng Zhang, Xinhui Kang, Siyu Wang, Peng Li, Ya-Qin Zhang, Weizhi Ma and Yang Liu · 2025
Closest in time.
“Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG”, 2025
Aditi Singh, Abul Ehtesham, Saket Kumar and Tala Khoei · 2025
Closest in time.
“DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning”, 2025
DeepSeek-AI et al · 2025
Closest in time.
“Guidelines for patient-centered documentation in the era of open notes: Qualitative study”
Anita Vanka, Katherine Johnston, Tom Delbanco, Catherine DesRoches, Annalays Garcia, Liz Salmi and Charlotte Blease · 2025
Closest in time.
“It’s time to bench the medical exam benchmark”
Inioluwa Raji, Roxana Daneshjou and Emily Alsentzer · 2025
Closest in time.
“Llm evaluators recognize and favor their own generations”
Arjun Panickssery, Samuel Bowman and Shi Feng · 2025
Closest in time.
“Hughes Hallucination Evaluation Model (HHEM) leaderboard” accessed: 2025-03-04, 2028
Huggingface · 2028
Closest in time.