Fetching the paper…
Reading the bibliography…
As large language models (LLMs) are increasingly deployed in healthcare, ensuring their safety, particularly within collaborative multi-agent configurations, is paramount.
The hare pcl-r: Some issues concerning its use and misuse
Robert D Hare · 1998
Earlier work this paper cites.
Principles of medical ethics, 2001
American Medical Association · 2001
Earlier work this paper cites.
The dark triad of personality: Narcissism, machiavellianism, and psychopathy
Delroy L Paulhus and Kevin M Williams · 2002
Earlier work this paper cites.
Introducing the short dark triad (sd3) a brief measure of dark personality traits
Daniel N Jones and Delroy L Paulhus · 2014
Earlier work this paper cites.
Pubmedqa: A dataset for biomedical research question answering
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W Cohen, and Xinghua Lu · 2019
Earlier work this paper cites.
Learning adversarial attack policies through multi-objective reinforcement learning
Javier García, Rubén Majadas, and Fernando Fernández · 2020
Earlier work this paper cites.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits · 2021
Earlier work this paper cites.
Meditron-70b: Scaling medical pretraining for large language models
Zeming Chen, Alejandro Hernández Cano, Angelika Romanou, Antoine Bonnet, Kyle Matoba, Francesco Salvi, Matteo Pagliardini, Simin Fan, Andreas Köpf, Amirkeivan Mohtashami, et al · 2023
Earlier work this paper cites.
The utility of chatgpt as an example of large language models in healthcare education, research and practice: Systematic review on the future perspectives and potential limitations
Malik Sallam · 2023
Earlier work this paper cites.
Agentbench: Evaluating llms as agents
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al · 2023
Earlier work this paper cites.
Camel: Communicative agents for" mind" exploration of large language model society
Guohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem · 2023
Earlier work this paper cites.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Earlier work this paper cites.
Attacking cooperative multi-agent reinforcement learning by adversarial minority influence
Simin Li, Jun Guo, Jingqiao Xiu, Yuwei Zheng, Pu Feng, Xin Yu, Aishan Liu, Yaodong Yang, Bo An, Wenjun Wu, et al · 2023
Earlier work this paper cites.
Chain-of-verification reduces hallucination in large language models
Shehzaad Dhuliawala, Mojtaba Komeili, Jing Xu, Roberta Raileanu, Xian Li, Asli Celikyilmaz, and Jason Weston · 2023
Earlier work this paper cites.
A knowledge-enhanced hierarchical reinforcement learning-based dialogue system for automatic disease diagnosis
Ying Zhu, Yameng Li, Yuan Cui, Tianbao Zhang, Daling Wang, Yifei Zhang, and Shi Feng · 2023
Earlier work this paper cites.
Cooperative dual medical ontology representation learning for clinical assisted decision-making
Muhao Xu, Zhenfeng Zhu, Youru Li, Shuai Zheng, Linfeng Li, Haiyan Wu, and Yao Zhao · 2023
Earlier work this paper cites.
Self-instruct: Aligning language models with self-generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2023
Earlier work this paper cites.
Prompt2model: Generating deployable models from natural language instructions
Vijay Viswanathan, Chenyang Zhao, Amanda Bertsch, Tongshuang Wu, and Graham Neubig · 2023
Earlier work this paper cites.
Can generalist foundation models outcompete special-purpose tuning? case study in medicine
Harsha Nori, Yin Tat Lee, Sheng Zhang, Dean Carignan, Richard Edgar, Nicolo Fusi, Nicholas King, Jonathan Larson, Yuanzhi Li, Weishung Liu, et al · 2023
Earlier work this paper cites.
Agent hospital: A simulacrum of hospital with evolvable medical agents
Junkai Li, Siyu Wang, Meng Zhang, Weitao Li, Yunghwei Lai, Xinhui Kang, Weizhi Ma, and Yang Liu · 2024
Earlier work this paper cites.
Agentic llm workflows for generating patient-friendly medical reports
Malavikha Sudarshan, Sophie Shih, Estella Yee, Alina Yang, John Zou, Cathy Chen, Quan Zhou, Leon Chen, Chinmay Singhal, and George Shih · 2024
Earlier work this paper cites.
Adaptive reasoning and acting in medical language agents
Abhishek Dutta and Yen-Che Hsiao · 2024
Earlier work this paper cites.
Mitigating cognitive biases in clinical decision-making through multi-agent conversations using large language models: simulation study
Yuhe Ke, Rui Yang, Sui An Lie, Taylor Xin Yi Lim, Yilin Ning, Irene Li, Hairil Rizal Abdullah, Daniel Shu Wei Ting, and Nan Liu · 2024
Earlier work this paper cites.
Medagents: Large language models as collaborators for zero-shot medical reasoning
Xiangru Tang, Anni Zou, Zhuosheng Zhang, Ziming Li, Yilun Zhao, Xingyao Zhang, Arman Cohan, and Mark Gerstein · 2024
Cited alongside, same era.
Triageagent: Towards better multi-agents collaborations for large language model-based clinical triage
Meng Lu, Brandon Ho, Dennis Ren, and Xuan Wang · 2024
Cited alongside, same era.
Mdagents: An adaptive collaboration of llms for medical decision-making
Yubin Kim, Chanwoo Park, Hyewon Jeong, Yik Siu Chan, Xuhai Xu, Daniel McDuff, Hyeonhoon Lee, Marzyeh Ghassemi, Cynthia Breazeal, and Hae Won Park · 2024
Cited alongside, same era.
Rareagents: Autonomous multi-disciplinary team for rare disease diagnosis and treatment
Xuanzhong Chen, Ye Jin, Xiaohao Mao, Lun Wang, Shuyang Zhang, and Ting Chen · 2024
Cited alongside, same era.
Medsafetybench: Evaluating and improving the medical safety of large language models
Ontomedrec: Logically-pretrained model-agnostic ontology encoders for medication recommendation
Weicong Tan, Weiqing Wang, Xin Zhou, Wray Buntine, Gordon Bingham, and Hongzhi Yin · 2024
Later among the works it cites.
Alignment of large language models in solving medical ethical dilemmas
Vera Sorin, Benjamin S Glicksberg, Panagiotis Korfiatis, Jeremy D Collins, Mei-Ean E Yeow, Megan Brandeland, Girish N Nadkarni, and Eyal Klang · 2024
Later among the works it cites.
Multi-expert prompting improves reliability, safety and usefulness of large language models
Do Long, Duong Yen, Luu Anh Tuan, Kenji Kawaguchi, Min-Yen Kan, and Nancy Chen · 2024
Later among the works it cites.
Metagpt: Meta programming for a multi-agent collaborative framework
Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, et al · 2024
Later among the works it cites.
Chatdev: Communicative agents for software development
Chen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tessa Han, Aounon Kumar, Chirag Agarwal, and Himabindu Lakkaraju · 2024
Cited alongside, same era.
Medfuzz: Exploring the robustness of large language models in medical question answering
Robert Osazuwa Ness, Katie Matton, Hayden Helm, Sheng Zhang, Junaid Bajwa, Carey E Priebe, and Eric Horvitz · 2024
Cited alongside, same era.
Large language models in healthcare and medical domain: A review
Zabir Al Nazi and Wei Peng · 2024
Cited alongside, same era.
Dimitrios P Panagoulias, Persephone Papatheodosiou, Anastasios P Palamidas, Mattheos Sanoudos, Evridiki Tsoureli-Nikita, Maria Virvou, and George A Tsihrintzis · 2024
Cited alongside, same era.
Psysafe: A comprehensive framework for psychological-based attack, defense, and evaluation of multi-agent system safety
Zaibin Zhang, Yongting Zhang, Lijun Li, Jing Shao, Hongzhi Gao, Yu Qiao, Lijun Wang, Huchuan Lu, and Feng Zhao · 2024
Cited alongside, same era.
Agentclinic: a multimodal agent benchmark to evaluate ai in simulated clinical environments
Samuel Schmidgall, Rojin Ziaei, Carl Harris, Eduardo Reis, Jeffrey Jopling, and Michael Moor · 2024
Cited alongside, same era.
Air-bench 2024: A safety benchmark based on risk categories from regulations and policies
Yi Zeng, Yu Yang, Andy Zhou, Jeffrey Ziwei Tan, Yuheng Tu, Yifan Mai, Kevin Klyman, Minzhou Pan, Ruoxi Jia, Dawn Song, et al · 2024
Cited alongside, same era.
R-judge: Benchmarking safety risk awareness for llm agents
Tongxin Yuan, Zhiwei He, Lingzhong Dong, Yiming Wang, Ruijie Zhao, Tian Xia, Lizhen Xu, Binglin Zhou, Fangqi Li, Zhuosheng Zhang, et al · 2024
Cited alongside, same era.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Later among the works it cites.
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al · 2024
Later among the works it cites.
Llama-3-meditron: An open-weight suite of medical llms based on llama-3.1
Alexandre Sallinen, Antoni-Joan Solergibert, Michael Zhang, Guillaume Boyé, Maud Dupont-Roc, Xavier Theimer-Lienhard, Etienne Boisson, Bastien Bernath, Hichem Hadhri, Antoine Tran, et al · 2025
Closest in time.
Pharmagents: Building a virtual pharma with large language model agents
Bowen Gao, Yanwen Huang, Yiqiao Liu, Wenxuan Xie, Wei-Ying Ma, Ya-Qin Zhang, and Yanyan Lan · 2025
Closest in time.
Kai Chen, Xinfeng Li, Tianpei Yang, Hewei Wang, Wei Dong, and Yang Gao · 2025
Closest in time.
A survey of llm-based agents in medicine: How far are we from baymax?
Wenxuan Wang, Zizhan Ma, Zheng Wang, Chenghan Wu, Wenting Chen, Xiang Li, and Yixuan Yuan · 2025
Closest in time.
Towards evaluating and building versatile large language models for medicine
Chaoyi Wu, Pengcheng Qiu, Jinxin Liu, Hongfei Gu, Na Li, Ya Zhang, Yanfeng Wang, and Weidi Xie · 2025
Closest in time.
Medagentsbench: Benchmarking thinking models and agent frameworks for complex medical reasoning
Xiangru Tang, Daniel Shao, Jiwoong Sohn, Jiapeng Chen, Jiayi Zhang, Jinyu Xiang, Fang Wu, Yilun Zhao, Chenglin Wu, Wenqi Shi, et al · 2025
Closest in time.
Medagentbench: Dataset for benchmarking llms as agents in medical applications
Yixing Jiang, Kameron C Black, Gloria Geng, Danny Park, Andrew Y Ng, and Jonathan H Chen · 2025
Closest in time.
Ailuminate: Introducing v1. 0 of the ai risk and reliability benchmark from mlcommons
Shaona Ghosh, Heather Frase, Adina Williams, Sarah Luger, Paul Röttger, Fazl Barez, Sean McGregor, Kenneth Fricklas, Mala Kumar, Kurt Bollacker, et al · 2025
Closest in time.
Toward expert-level medical question answering with large language models
Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Mohamed Amin, Le Hou, Kevin Clark, Stephen R Pfohl, Heather Cole-Lewis, et al · 2025
Closest in time.
Vital: A new dataset for benchmarking pluralistic alignment in healthcare
Anudeex Shetty, Amin Beheshti, Mark Dras, and Usman Naseem · 2025
Closest in time.
Prompt injection detection and mitigation via ai multi-agent nlp frameworks
Diego Gosmar, Deborah A Dahl, and Dario Gosmar · 2025
Closest in time.
Emerging cyber attack risks of medical ai agents
Jianing Qiu, Lin Li, Jiankai Sun, Hao Wei, Zhe Xu, Kyle Lam, and Wu Yuan · 2025
Closest in time.
Red-teaming llm multi-agent systems via communication attacks
Pengfei He, Yupin Lin, Shen Dong, Han Xu, Yue Xing, and Hui Liu · 2025
Closest in time.
Medical mllm is vulnerable: Cross-modality jailbreak and mismatched attacks on medical multimodal large language models
Xijie Huang, Xinyuan Wang, Hantao Zhang, Yinghao Zhu, Jiawen Xi, Jingkun An, Hao Wang, Hao Liang, and Chengwei Pan · 2025
Closest in time.
Hierarchical divide-and-conquer for fine-grained alignment in llm-based medical evaluation
Shunfan Zheng, Xiechi Zhang, Gerard de Melo, Xiaoling Wang, and Linlin Wang · 2025
Closest in time.
M3hf: Multi-agent reinforcement learning from multi-phase human feedback of mixed quality
Ziyan Wang, Zhicheng Zhang, Fei Fang, and Yali Du · 2025
Closest in time.