Fetching the paper…
Reading the bibliography…
Multimodal/vision language models (VLMs) are increasingly being deployed in healthcare settings worldwide, necessitating robust benchmarks to ensure their safety, efficacy, and fairness.
Large language models in medicine
Arun James Thirunavukarasu, Darren Shu Jeng Ting, Kabilan Elangovan, Laura Gutierrez, Ting Fang Tan, and Daniel Shu Wei Ting. 2023 · 1940
Earlier work this paper cites.
Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit
Jacob Cohen. 1968 · 1968
Earlier work this paper cites.
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2020 · 2009
Earlier work this paper cites.
MLEC-QA: A Chinese Multi-Choice Biomedical Question Answering Dataset
Jing Li, Shangping Zhong, and Kaizhi Chen. 2021a · 2021
Earlier work this paper cites.
MLEC-QA: A Chinese Multi-Choice Biomedical Question Answering Dataset
Jing Li, Shangping Zhong, and Kaizhi Chen. 2021b · 2021
Earlier work this paper cites.
Ethical and social risks of harm from language models
Laura Weidinger et al. 2021 · 2021
Earlier work this paper cites.
The 116th National Medical Examination Questions and Answers
Japanese Ministry of Health Labour and Welfare 2022 · 2022
Earlier work this paper cites.
Medmcqa : A large-scale multi-subject multi-choice dataset for medical domain question answering
Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2022 · 2022
Earlier work this paper cites.
Generative ai for infectious diseases: An evaluation of chatgpt for medical translation
Dillon C Adam et al. 2023 · 2023
Earlier work this paper cites.
Comparing ChatGPT and GPT-4 performance in USMLE soft skill assessments
Dana Brin, Vera Sorin, Akhil Vaid, Ali Soroush, Benjamin S. Glicksberg, Alexander W. Charney, Girish Nadkarni, and Eyal Klang. 2023 · 2023
Earlier work this paper cites.
The future landscape of large language models in medicine
Jan Clusmann, Fiona R. Kolbinger, Hannah Sophie Muti, Zunamys I. Carrero, Jan-Niklas Eckardt, Narmin Ghaffari Laleh, Chiara Maria Lavinia Löffler, Sophie-Caroline Schwarzkopf, Michaela Unger, Gregory P. Veldhuizen, Sophia J. Wagner, and Jakob Nikolas Kather. 2023 · 2023
Earlier work this paper cites.
A bibliometric review of large language models research from 2017 to 2023
Lizhou Fan, Lingyao Li, Zihui Ma, Sanggyu Lee, Huizi Yu, and Libby Hemphill. 2023 · 2023
Earlier work this paper cites.
Performance of ChatGPT-4 in answering questions from the Brazilian National Examination for Medical Degree Revalidation
Mauro Gobira, Luis Filipe Nakayama, Rodrigo Moreira, Eric Andrade, Caio Vinicius Saito Regatieri, and Rubens Belfort Jr. 2023 · 2023
Cited alongside, same era.
The 117th National Medical Examination Questions and Answers
Japanese Ministry of Health Labour and Welfare 2023 · 2023
Cited alongside, same era.
Performance of chatgpt on usmle: Potential for ai-assisted medical education using large language models
T. H. Kung, M. Cheatham, A. Medenilla, C. Sillos, L. De Leon, C. Elepaño, M. Madriaga, R. Aggabao, G. Diaz-Candido, J. Maningo, and V. Tseng. 2023 · 2023
Cited alongside, same era.
Culturalvqa: A new frontier in vision and language understanding
Yiyi Wang et al. 2023 · 2023
Cited alongside, same era.
Chaoyi Wu, Jiayu Lei, Qiaoyu Zheng, Weike Zhao, Weixiong Lin, Xiaoman Zhang, Xiao Zhou, Ziheng Zhao, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2023 · 2023
MedQA-SWE - a clinical question & answer dataset for Swedish
Niclas Hertzberg and Anna Lokrantz. 2024b · 2024
Closest in time.
The 118th National Medical Examination Questions and Answers
Japanese Ministry of Health Labour and Welfare 2024 · 2024
Closest in time.
Sunjun Kweon, Byungjin Choi, Minkyu Kim, Rae Woong Park, and Edward Choi. 2024 · 2024
Closest in time.
Performance of chatgpt across different versions in medical licensing examinations worldwide: Systematic review and meta-analysis
M. Liu, T. Okuhara, X. Chang, R. Shirabe, Y. Nishiie, H. Okada, and T. Kiuchi. 2024 · 2024
Closest in time.
Seeing beyond borders: Evaluating llms in multilingual ophthalmological question answering
David Restrepo, Luis Filipe Nakayama, Robyn Gayle Dychiao, Chenwei Wu, Liam G. McCoy, Jose Carlo Artiaga, Marisa Cobanaj, João Matos, Jack Gallifant, Danielle S. Bitterman, Vincenz Ferrer, Yindalon Aphinyanaphongs, and Leo Anthony Celi. 2024a · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Multimodal chatgpt for medical applications: an experimental study of gpt-4v
Zhiling Yan, Kai Zhang, Rong Zhou, Lifang He, Xiang Li, and Lichao Sun. 2023 · 2023
Cited alongside, same era.
Foundation metrics for evaluating effectiveness of healthcare conversations powered by generative AI
Mahyar Abbasian, Elahe Khatibi, Iman Azimi, David Oniani, Zahra Shakeri Hossein Abad, Alexander Thieme, Ram Sriram, Zhongqi Yang, Yanshan Wang, Bryant Lin, Olivier Gevaert, Li-Jia Li, Ramesh Jain, and Amir M. Rahmani. 2024 · 2024
Cited alongside, same era.
Exploring the landscape of large language models in medical question answering
Andrew M. Bean, Karolina Korgul, Felix Krones, Robert McCraith, and Adam Mahdi. 2024 · 2024
Cited alongside, same era.
Vlmevalkit: An open-source toolkit for evaluating large multi-modality models
Haodong Duan, Junming Yang, Yuxuan Qiao, Xinyu Fang, Lin Chen, Yuan Liu, Xiaoyi Dong, Yuhang Zang, Pan Zhang, Jiaqi Wang, Dahua Lin, and Kai Chen. 2024 · 2024
Cited alongside, same era.
Language models are surprisingly fragile to drug names in biomedical benchmarks
Jack Gallifant, Shan Chen, Pedro Moreira, Nikolaj Munch, Mingye Gao, Jackson Pond, Leo Anthony Celi, Hugo Aerts, Thomas Hartvigsen, and Danielle Bitterman. 2024 · 2024
Cited alongside, same era.
MedQA-SWE - a clinical question & answer dataset for Swedish
Niclas Hertzberg and Anna Lokrantz. 2024a · 2024
Cited alongside, same era.
IMA - Israel Medicine Association
Israel Medicine Association
Cited in the paper.
Closest in time.
Unintended Impacts of LLM Alignment on Global Representation
Michael J. Ryan, William Held, and Diyi Yang. 2024 · 2024
Closest in time.
Capabilities of gemini models in medicine
Khaled Saab, Tao Tu, Wei-Hung Weng, Ryutaro Tanno, David Stutz, Ellery Wulczyn, Fan Zhang, Tim Strother, Chunjong Park, Elahe Vedadi, Juanma Zambrano Chaves, Szu-Yeu Hu, Mike Schaekermann, Aishwarya Kamath, Yong Cheng, David G. T. Barrett, Cathy Cheung, Basil Mustafa, Anil Palepu, Daniel McDuff, Le Hou, Tomer Golany, Luyang Liu, Jean baptiste Alayrac, Neil Houlsby, Nenad Tomasev, Jan Freyberg, Charles Lau, Jonas Kemp, Jeremy Lai, Shekoofeh Azizi, Kimberly Kanada, SiWai Man, Kavita Kulkarni, Ruoxi Sun, Siamak Shakeri, Luheng He, Ben Caine, Albert Webson, Natasha Latysheva, Melvin Johnson, Philip Mansfield, Jian Lu, Ehud Rivlin, Jesper Anderson, Bradley Green, Renee Wong, Jonathan Krause, Jonathon Shlens, Ewa Dominowska, S. M. Ali Eslami, Katherine Chou, Claire Cui, Oriol Vinyals, Koray Kavukcuoglu, James Manyika, Jeff Dean, Demis Hassabis, Yossi Matias, Dale Webster, Joelle Barral, Greg Corrado, Christopher Semturs, S. Sara Mahdavi, Juraj Gottweis, Alan Karthikesalingam, and Vivek Natarajan. 2024 · 2024
Closest in time.
Hugging Face releases a benchmark for testing generative AI on health tasks
Kyle Wiggers. 2024 · 2024
Closest in time.
Large language models in biomedical and health informatics: A bibliometric review
Huizi Yu, Lizhou Fan, Lingyao Li, Jiayan Zhou, Zihui Ma, Lu Xian, Wenyue Hua, Sijia He, Mingyu Jin, Yongfeng Zhang, Ashvin Gandhi, and Xin Ma. 2024 · 2024
Closest in time.
Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study
Travis Zack, Eric Lehman, Mirac Suzgun, Jorge A Rodriguez, Leo Anthony Celi, Judy Gichoya, Dan Jurafsky, Peter Szolovits, David W Bates, Raja-Elie E Abdulnour, Atul J Butte, and Emily Alsentzer. 2024 · 2024
Closest in time.
Aya model: An instruction finetuned open-access multilingual language model
Ahmet Üstün, Viraat Aryabumi, Zheng-Xin Yong, Wei-Yin Ko, Daniel D’souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, Freddie Vargus, Phil Blunsom, Shayne Longpre, Niklas Muennighoff, Marzieh Fadaee, Julia Kreutzer, and Sara Hooker. 2024 · 2024
Closest in time.