Fetching the paper…
Reading the bibliography…
Several studies showed that Large Language Models (LLMs) can answer medical questions correctly, even outperforming the average human score in some medical exams.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits · 2009
Earlier work this paper cites.
Frequency and types of patient-reported errors in electronic health record ambulatory care notes
Sigall K. Bell, Tom Delbanco, Joann G. Elmore, Patricia S. Fitzgerald, Alan Fossa, Kendall Harcourt, Suzanne G. Leveille, Thomas H. Payne, Rebecca A. Stametz, Jan Walker, and Catherine M. DesRoches · 2020
Earlier work this paper cites.
BLEURT: learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur P. Parikh · 2020
Earlier work this paper cites.
SemEval-2020 task 4: Commonsense validation and explanation
Cunxiang Wang, Shuailong Liang, Yili Jin, Yilong Wang, Xiaodan Zhu, and Yue Zhang · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi · 2020
Earlier work this paper cites.
CREAK: A dataset for commonsense reasoning over entity knowledge
Yasumasa Onoe, Michael J. Q. Zhang, Eunsol Choi, and Greg Durrett · 2021
Earlier work this paper cites.
BECEL: benchmark for consistency evaluation of language models
Myeongjun Jang, Deuk Sin Kwon, and Thomas Lukasiewicz · 2022
Earlier work this paper cites.
An investigation of evaluation metrics for automated medical note generation
Asma Ben Abacha, Wen-wai Yim, George Michalopoulos, and Thomas Lin · 2023
Cited alongside, same era.
How does chatgpt perform on the united states medical licensing examination (usmle)? the implications of large language models for medical education and knowledge assessment
Aidan Gilson, Conrad W Safranek, Thomas Huang, Vimig Socrates, Ling Chi, Richard Andrew Taylor, and David Chartas · 2023
Cited alongside, same era.
Consistency analysis of chatgpt
Myeongjun Jang and Thomas Lukasiewicz · 2023
Cited alongside, same era.
Assessing the accuracy and reliability of ai-generated medical responses: An evaluation of the chat-gpt model
Douglas Johnson, Rachel Goodman, J Patrinely, Cosby Stone, Eli Zimmerman, Rebecca Donald, Sam Chang, Sean Berkowitz, Avni Finn, Eiman Jahangir, Elizabeth Scoville, Tyler Reese, Debra Friedman, Julie Bastarache, Yuri van der Heijden, Jordan Wright, Nicholas Carter, Matthew Alexander, Jennifer Choe, Cody Chastain, John Zic, Sara Horst, Isik Turker, Rajiv Agarwal, Evan Osmundson, Kamran Idrees, Colleen Kieman, Chandrasekhar Padmanabhan, Christina Bailey, Cameron Schlegel, Lola Chambless, Mike Gibson, Travis Osterman, and Lee Wheles · 2023
Cited alongside, same era.
Phi-3 technical report: A highly capable language model locally on your phone
Marah I Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat S. Behl, Alon Benhaim, Misha Bilenko, Johan Bjorck, Sébastien Bubeck, Martin Cai, Caio César Teodoro Mendes, Weizhu Chen, Vishrav Chaudhary, Parul Chopra, Allie Del Giorno, Gustavo de Rosa, Matthew Dixon, Ronen Eldan, Dan Iter, Amit Garg, Abhishek Goswami, Suriya Gunasekar, Emman Haider, Junheng Hao, Russell J. Hewett, Jamie Huynh, Mojan Javaheripi, Xin Jin, Piero Kauffmann, Nikos Karampatziakis, Dongwoo Kim, Mahoud Khademi, Lev Kurilenko, James R. Lee, Yin Tat Lee, Yuanzhi Li, Chen Liang, Weishung Liu, Eric Lin, Zeqi Lin, Piyush Madan, Arindam Mitra, Hardik Modi, Anh Nguyen, Brandon Norick, Barun Patra, Daniel Perez-Becker, Thomas Portet, Reid Pryzant, Heyang Qin, Marko Radmilac, Corby Rosset, Sambudha Roy, Olatunji Ruwase, Olli Saarikivi, Amin Saied, Adil Salim, Michael Santacroce, Shital Shah, Ning Shang, Hiteshi Sharma, Xia Song, Masahiro Tanaka, Xin Wang, Rachel Ward, Guanhua Wang, Philipp Witte, Michael Wyatt, Can Xu, Jiahang Xu, Sonali Yadav, Fan Yang, Ziyi Yang, Donghan Yu, Chengruidong Zhang, Cyril Zhang, Jianwen Zhang, Li Lyna Zhang, Yi Zhang, Yue Zhang, Yunan Zhang, and Xiren Zhou · 2024
Closest in time.
Claude 3.5 sonnet
Anthropic · 2024
Closest in time.
Overview of the mediqa-corr 2024 shared task on medical error detection and correction
Asma Ben Abacha, Wen wai Yim, Yujuan Fu, Zhaoyi Sun, Fei Xia, and Meliha Yetisgen · 2024
Closest in time.
The effect of using a large language model to respond to patient messages
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Embracing large language models for medical applications: Opportunities and challenges
Mert Karabacak and Konstantinos Margetis · 2023
Cited alongside, same era.
Performance of Large Language Models on a Neurology Board–Style Examination
Marc Cicero Schubert, Wolfgang Wick, and Varun Venkataramani · 2023
Cited alongside, same era.
Towards expert-level medical question answering with large language models
Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Le Hou, Kevin Clark, Stephen Pfohl, Heather Cole-Lewis, Darlene Neal, Mike Schaekermann, Amy Wang, Mohamed Amin, Sami Lachgar, Philip Andrew Mansfield, Sushant Prakash, Bradley Green, Ewa Dominowska, Blaise Agüera y Arcas, Nenad Tomasev, Yun Liu, Renee Wong, Christopher Semturs, S. Sara Mahdavi, Joelle K. Barral, Dale R. Webster, Gregory S. Corrado, Yossi Matias, Shekoofeh Azizi, Alan Karthikesalingam, and Vivek Natarajan · 2023
Cited alongside, same era.
Evaluating large language models on medical evidence summarization
Liyan Tang, Zhaoyi Sun, Betina Ross S. Idnay, Jordan G. Nestor, Ali Soroush, Pierre A. Elias, Ziyang Xu, Ying Ding, Greg Durrett, Justin F. Rousseau, Chunhua Weng, and Yifan Peng · 2023
Cited alongside, same era.
Shan Chen, Marco Guevara, Shalini Moningi, Frank Hoebers, Hesham Elhalawani, Benjamin H Kann, Fallon E Chipidza, Jonathan Leeman, Hugo J W L Aerts, Timothy Miller, Guergana K Savova, Jack Gallifant, Leo A Celi, Raymond H Mak, Maryam Lustberg, Majid Afshar, and Danielle S Bitterman · 2024
Closest in time.
Gemini 2.0 flash
Google · 2024
Closest in time.
gpt-3.5-turbo
OpenAI · 2024
Closest in time.
Gpt-4o mini
OpenAI · 2024
Closest in time.
Diagnostic reasoning prompts reveal the potential for large language model interpretability in medicine
Thomas Savage, Ashwin Nayak, Robert Gallo, Ekanath Rangan, and Jonathan H. Chen · 2024
Closest in time.