Fetching the paper…
Reading the bibliography…
Large language models (LLMs) deployed as agents introduce significant safety risks in clinical settings due to their potential for error and single points of failure.
Multi-criteria clinical decision support: a primer on the use of multiple-criteria decision-making methods to promote evidence-based, patient-centered healthcare
James G Dolan · 2010
Earlier work this paper cites.
An ethical hierarchy for decision making during medical emergencies
Patrick D Lyden, Brett C Meyer, Thomas M Hemmen, and Karen S Rapp · 2010
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Geoffrey Irving, Paul Christiano, and Dario Amodei · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg · 2018
Earlier work this paper cites.
Pubmedqa: A dataset for biomedical research question answering
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu · 2019
Earlier work this paper cites.
Role of artificial intelligence in patient safety outcomes: Systematic literature review
Avishek Choudhury and Onur Asan · 2020
Earlier work this paper cites.
To what extent does hierarchical leadership affect health care outcomes?
Navindi Fernandopulle · 2021
Earlier work this paper cites.
Domain-specific language model pretraining for biomedical natural language processing
Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon · 2021
Earlier work this paper cites.
Unsolved problems in ml safety
Dan Hendrycks, Nicholas Carlini, John Schulman, and Jacob Steinhardt · 2021
Earlier work this paper cites.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits · 2021
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al · 2022
Earlier work this paper cites.
Measuring progress on scalable oversight for large language models
Samuel R Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamilė Lukošiūtė, Amanda Askell, Andy Jones, Anna Chen, et al · 2022
Earlier work this paper cites.
Measuring progress on scalable oversight for large language models, 2022
SR Bowman, J Hyun, E Perez, E Chen, C Pettit, S Heiner, K Lukošiute, A Askell, A Jones, A Chen, et al · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al · 2022
Earlier work this paper cites.
Inner monologue: Embodied reasoning through planning with language models
Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Improving factuality and reasoning in language models through multiagent debate, 2023
Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch · 2023
Earlier work this paper cites.
Improving language model negotiation with self-play and in-context learning from ai feedback
Yao Fu, Hao Peng, Tushar Khot, and Mirella Lapata · 2023
Earlier work this paper cites.
The capability of large language models to measure psychiatric functioning
Isaac R Galatzer-Levy, Daniel McDuff, Vivek Natarajan, Alan Karthikesalingam, and Matteo Malgaroli · 2023
Earlier work this paper cites.
Interprofessional collaboration in complex patient care transition: a qualitative multi-perspective analysis
Franziska Geese and Kai-Uwe Schmitt · 2023
Earlier work this paper cites.
Reasoning with language model is planning with world model
Shibo Hao, Yi Gu, Haodi Ma, Joshua Hong, Zhen Wang, Daisy Wang, and Zhiting Hu · 2023
Earlier work this paper cites.
Performance of chatgpt on usmle: Potential for ai-assisted medical education using large language models
Tiffany H. Kung, Morgan Cheatham, Arielle Medenilla, Czarina Sillos, Lorie De Leon, Camille Elepaño, Maria Madriaga, Rimel Aggabao, Giezel Diaz-Candido, James Maningo, and Victor Tseng · 2023
Earlier work this paper cites.
Camel: Communicative agents for "mind" exploration of large language model society, 2023
Guohao Li, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem · 2023
Earlier work this paper cites.
Encouraging divergent thinking in large language models through multi-agent debate, 2023
Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Zhaopeng Tu, and Shuming Shi · 2023
Earlier work this paper cites.
Can large language models reason about medical questions?, 2023
Valentin Liévin, Christoffer Egeberg Hother, Andreas Geert Motzfeldt, and Ole Winther · 2023
Earlier work this paper cites.
Towards accurate differential diagnosis with large language models
Daniel McDuff, Mike Schaekermann, Tao Tu, Anil Palepu, Amy Wang, Jake Garrison, Karan Singhal, Yash Sharma, Shekoofeh Azizi, Kavita Kulkarni, et al · 2023
Earlier work this paper cites.
Foundation models for generalist medical artificial intelligence
Michael Moor, Oishi Banerjee, Zahra Shakeri Hossein Abad, Harlan M. Krumholz, Jure Leskovec, Eric J. Topol, and Pranav Rajpurkar · 2023
Earlier work this paper cites.
Practices for governing agentic ai systems
OpenAI · 2023
Earlier work this paper cites.
Med-halt: Medical domain hallucination test for large language models
Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu · 2023
Earlier work this paper cites.
Practices for governing agentic ai systems
Yonadav Shavit, Sandhini Agarwal, Miles Brundage, Steven Adler, Cullen O’Keefe, Rosie Campbell, Teddy Lee, Pamela Mishkin, Tyna Eloundou, Alan Hickey, et al · 2023
Earlier work this paper cites.
Hugginggpt: Solving AI tasks with chatgpt and its friends in huggingface
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang · 2023
Earlier work this paper cites.
Autogpt, 2023
Significant Gravitas · 2023
Earlier work this paper cites.
Large language models encode clinical knowledge
Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry Payne, Martin Seneviratne, Paul Gamble, Chris Kelly, Abubakr Babiker, Nathanael Schärli, Aakanksha Chowdhery, Philip Mansfield, Dina Demner-Fushman, Blaise Agüera y Arcas, Dale Webster, Greg S. Corrado, Yossi Matias, Katherine Chou, Juraj Gottweis, Nenad Tomasev, Yun Liu, Alvin Rajkomar, Joelle Barral, Christopher Semturs, Alan Karthikesalingam, and Vivek Natarajan · 2023
Earlier work this paper cites.
Medagents: Large language models as collaborators for zero-shot medical reasoning
Xiangru Tang, Anni Zou, Zhuosheng Zhang, Ziming Li, Yilun Zhao, Xingyao Zhang, Arman Cohan, and Mark Gerstein · 2023
Earlier work this paper cites.
Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change
Karthik Valmeekam, Matthew Marquez, Alberto Olmo, Sarath Sreedharan, and Subbarao Kambhampati · 2023
Earlier work this paper cites.
Autogen: Enabling next-gen llm applications via multi-agent conversation, 2023
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W White, Doug Burger, and Chi Wang · 2023
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao · 2023
Cited alongside, same era.
Almanac: Retrieval-augmented language models for clinical medicine, 2023
Cyril Zakka, Akash Chaurasia, Rohan Shad, Alex R. Dalal, Jennifer L. Kim, Michael Moor, Kevin Alexander, Euan Ashley, Jack Boyd, Kathleen Boyd, Karen Hirsch, Curt Langlotz, Joanna Nelson, and William Hiesinger · 2023
Cited alongside, same era.
Safetybench: Evaluating the safety of large language models
Zhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun, Yongkang Huang, Chong Long, Xiao Liu, Xuanyu Lei, Jie Tang, and Minlie Huang · 2023
Cited alongside, same era.
Context-aware medical systems within healthcare environments: A systematic scoping review to identify subdomains and significant medical contexts
Cognitive architectures for language agents
Theodore Sumers, Shunyu Yao, Karthik Narasimhan, and Thomas Griffiths · 2024
Later among the works it cites.
Large language models seem miraculous, but science abhors miracles, 2024
Peter Szolovits · 2024
Later among the works it cites.
Medagents: Large language models as collaborators for zero-shot medical reasoning, 2024
Xiangru Tang, Anni Zou, Zhuosheng Zhang, Ziming Li, Yilun Zhao, Xingyao Zhang, Arman Cohan, and Mark Gerstein · 2024
Later among the works it cites.
Magis: Llm-based multi-agent framework for github issue resolution
Wei Tao, Yucheng Zhou, Yanlin Wang, Wenqiang Zhang, Hongyu Zhang, and Yu Cheng · 2024
Later among the works it cites.
Adapted large language models can outperform medical experts in clinical text summarization
Dave Van Veen, Cara Van Uden, Louis Blankemeier, Jean-Benoit Delbrouck, Asad Aali, Christian Bluethgen, Anuj Pareek, Malgorzata Polacin, Eduardo Pontes Reis, Anna Seehofnerová, Nidhi Rohatgi, Poonam Hosamani, William Collins, Neera Ahuja, Curtis P. Langlotz, Jason Hom, Sergios Gatidis, John Pauly, and Akshay S. Chaudhari · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Michael Zon, Guha Ganesh, M Jamal Deen, and Qiyin Fang · 2023
Cited alongside, same era.
Medhallbench: A new benchmark for assessing hallucination in medical large language models
Kaiwen Zuo and Yirui Jiang · 2023
Cited alongside, same era.
Medhalu: Hallucinations in responses to healthcare queries by large language models
Vibhor Agarwal, Yiqiao Jin, Mohit Chandra, Munmun De Choudhury, Srijan Kumar, and Nishanth Sastry · 2024
Cited alongside, same era.
A survey on llm-based agentic workflows and llm-profiled components
Anonymous · 2024
Cited alongside, same era.
Building effective ai agents
Anthropic · 2024
Cited alongside, same era.
Diverse and effective red teaming with auto-generated rewards and multi-step reinforcement learning
Alex Beutel, Kai Xiao, Johannes Heidecke, and Lilian Weng · 2024
Cited alongside, same era.
Workarena++: Towards compositional planning and reasoning-based common knowledge work tasks
Léo Boisvert, Megh Thakkar, Maxime Gasse, Massimo Caccia, Thibault Le Sellier De Chezelles, Quentin Cappart, Nicolas Chapados, Alexandre Lacoste, and Alexandre Drouin · 2024
Cited alongside, same era.
Red teaming large language models in medicine: real-world insights on model behavior
Crystal T Chang, Hodan Farah, Haiwen Gui, Shawheen Justin Rezaei, Charbel Bou-Khalil, Ye-Jean Park, Akshay Swaminathan, Jesutofunmi A Omiye, Akaash Kolluri, Akash Chaurasia, et al · 2024
Cited alongside, same era.
Autopatent: A multi-agent framework for automatic patent generation
Qiyao Wang, Shiwen Ni, Huaren Liu, Shule Lu, Guhong Chen, Xi Feng, Chi Wei, Qiang Qu, Hamid Alinejad-Rokny, Yuan Lin, et al · 2024
Later among the works it cites.
Clinicallab: Aligning agents for multi-departmental clinical diagnostics in the real world
Weixiang Yan, Haitian Liu, Tengxiao Wu, Qian Chen, Wen Wang, Haoyuan Chai, Jiayi Wang, Weishan Zhao, Yixin Zhang, Renjun Zhang, et al · 2024
Later among the works it cites.
Autodefense: Multi-agent llm defense against jailbreak attacks
Yifan Zeng, Yiran Wu, Xiao Zhang, Huazheng Wang, and Qingyun Wu · 2024
Later among the works it cites.
Longagent: Scaling language models to 128k context through multi-agent collaboration
Jun Zhao, Can Zu, Hao Xu, Yi Lu, Wei He, Yiwen Ding, Tao Gui, Qi Zhang, and Xuanjing Huang · 2024
Later among the works it cites.
On prompt-driven safeguarding for large language models
Chujie Zheng, Fan Yin, Hao Zhou, Fandong Meng, Jie Zhou, Kai-Wei Chang, Minlie Huang, and Nanyun Peng · 2024
Later among the works it cites.
International Journal of Scientific Research in Computer Science, Engineering and Information Technology , 11:567–575, 03 2025
Agentic workflows in healthcare: Advancing clinical efficiency through ai integration · 2025
Closest in time.
Survey on evaluation of llm-based agents
Anonymous · 2025
Closest in time.
Scaling enterprise ai in healthcare: the role of governance in risk mitigation frameworks
Andreea Bodnari and John Travis · 2025
Closest in time.
Agents or workflows?
Louis Bouchard · 2025
Closest in time.
Medsentry: Understanding and mitigating safety risks in medical llm multi-agent systems
Kai Chen, Taihang Zhen, Hewei Wang, Kailai Liu, Xinfeng Li, Jing Huo, Tianpei Yang, Jinfeng Xu, Wei Dong, and Yang Gao · 2025
Closest in time.
Differences in technical and clinical perspectives on ai validation in cancer imaging: mind the gap!
Ioanna Chouvarda, Sara Colantonio, Ana SC Verde, Ana Jimenez-Pastor, Leonor Cerdá-Alberich, Yannick Metz, Lithin Zacharias, Shereen Nabhani-Gebara, Maciej Bobowicz, Gianna Tsakou, et al · 2025
Closest in time.
Bridging the gap: From ai success in clinical trials to real-world healthcare implementation—a narrative review
Rabie Adel El Arab, Mohammad S Abu-Mahfouz, Fuad H Abuadas, Husam Alzghoul, Mohammed Almari, Ahmad Ghannam, and Mohamed Mahmoud Seweid · 2025
Closest in time.
Scaling laws for scalable oversight
Joshua Engels, David D Baek, Subhash Kantamneni, and Max Tegmark · 2025
Closest in time.
Cost-of-pass: An economic framework for evaluating language models
Mehmet Hamza Erol, Batu El, Mirac Suzgun, Mert Yuksekgonul, and James Zou · 2025
Closest in time.
Developing agentic ai workflows with safety and accuracy
Fiddler AI · 2025
Closest in time.
Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, et al · 2025
Closest in time.
The anatomy of a personal health agent
A Ali Heydari, Ken Gu, Vidya Srinivas, Hong Yu, Zhihan Zhang, Yuwei Zhang, Akshay Paruchuri, Qian He, Hamid Palangi, Nova Hammerquist, et al · 2025
Closest in time.
Next-generation agentic ai for transforming healthcare
Nalan Karunanayake · 2025
Closest in time.
Healthcare Agents: Large Language Models in Health Prediction and Decision-Making
Yubin Kim · 2025
Closest in time.
Medical hallucination in foundation models and their impact on healthcare
Yubin Kim, Hyewon Jeong, Shen Chen, Shuyue Stella Li, Mingyu Lu, Kumail Alhamoud, Jimin Mun, Cristina Grau, Minseok Jung, Rodrigo R Gameiro, et al · 2025
Closest in time.
A black swan hypothesis: The role of human irrationality in ai safety
Hyunin Lee, Chanwoo Park, David Abel, and Ming Jin · 2025
Closest in time.
Ai agents vs. chatbots, workflows, gpts: A guide to ai paradigms
Mindset.ai · 2025
Closest in time.
Towards a hipaa compliant agentic ai system in healthcare
Subash Neupane, Shaswata Mitra, Sudip Mittal, and Shahram Rahimi · 2025
Closest in time.
Flow: Modularized agentic workflow automation
Boye Niu, Yiliao Song, Kai Lian, Yifan Shen, Yu Yao, Kun Zhang, and Tongliang Liu · 2025
Closest in time.
Towards conversational ai for disease management
Anil Palepu, Valentin Liévin, Wei-Hung Weng, Khaled Saab, David Stutz, Yong Cheng, Kavita Kulkarni, S Sara Mahdavi, Joëlle Barral, Dale R Webster, et al · 2025
Closest in time.
Andrew Sellergren, Sahar Kazemzadeh, Tiam Jaroensri, Atilla Kiraly, Madeleine Traverse, Timo Kohlberger, Shawn Xu, Fayaz Jamil, Cían Hughes, Charles Lau, et al · 2025
Closest in time.
Toward expert-level medical question answering with large language models
Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Mohamed Amin, Le Hou, Kevin Clark, Stephen R Pfohl, Heather Cole-Lewis, et al · 2025
Closest in time.
What are agentic workflows? patterns, use cases, examples, and challenges
Weaviate · 2025
Closest in time.
The ai scientist-v2: Workshop-level automated scientific discovery via agentic tree search
Yutaro Yamada, Robert Tjarko Lange, Cong Lu, Shengran Hu, Chris Lu, Jakob Foerster, Jeff Clune, and David Ha · 2025
Closest in time.
The rise of agentic ai teammates in medicine
James Zou and Eric J Topol · 2025
Closest in time.