Fetching the paper…
Reading the bibliography…
Recent advancements in large language model(LLM) performance on medical multiple choice question (MCQ) benchmarks have stimulated interest from healthcare providers and patients globally.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Xiaodong Song, and Jacob Steinhardt. 2020 · 2009
Earlier work this paper cites.
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2020 · 2009
Earlier work this paper cites.
Bridging the gap between consumers’ medication questions and trusted answers
Asma Ben Abacha, Yassine Mrabet, Mark E. Sharp, Travis R. Goodwin, Sonya E. Shooshan, and Dina Demner-Fushman. 2019 · 2019
Earlier work this paper cites.
PubMedQA: A dataset for biomedical research question answering
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. 2019 · 2019
Earlier work this paper cites.
LiveQA: A question answering dataset over sports live
Qianying Liu, Sicong Jiang, Yizhong Wang, and Sujian Li. 2020 · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Earlier work this paper cites.
A fine-grained analysis of bertscore
Michael Hanna and Ondřej Bojar. 2021 · 2021
Earlier work this paper cites.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021 · 2021
Earlier work this paper cites.
Q-pain: A question answering dataset to measure social bias in pain management
C’ecile Log’e, Emily L. Ross, David Yaw Amoah Dadey, Saahil Jain, Adriel Saporta, Andrew Y. Ng, and Pranav Rajpurkar. 2021 · 2021
Earlier work this paper cites.
Questeval: Summarization asks for fact-based evaluation
Thomas Scialom, Paul-Alexis Dray, Patrick Gallinari, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, and Alex Wang. 2021 · 2021
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Earlier work this paper cites.
Meditron-70b: Scaling medical pretraining for large language models
Zeming Chen, Alejandro Hernández Cano, Angelika Romanou, Antoine Bonnet, Kyle Matoba, Francesco Salvi, Matteo Pagliardini, Simin Fan, Andreas Köpf, Amirkeivan Mohtashami, et al. 2023 · 2023
Cited alongside, same era.
Use of gpt-4 to diagnose complex clinical cases
Alexander V Eriksen, Sören Möller, and Jesper Ryg. 2023 · 2023
Cited alongside, same era.
Llms: A promising new tool for improving healthcare in low-resource nations
Agasthya Gangavarapu. 2023 · 2023
Cited alongside, same era.
Bioasq-qa: A manually curated corpus for biomedical question answering
Anastasia Krithara, Anastasios Nentidis, Konstantinos Bougiatiotis, and Georgios Paliouras. 2023 · 2023
Cited alongside, same era.
Performance of chatgpt on usmle: potential for ai-assisted medical education using large language models
Tiffany H Kung, Morgan Cheatham, Arielle Medenilla, Czarina Sillos, Lorie De Leon, Camille Elepaño, Maria Madriaga, Rimel Aggabao, Giezel Diaz-Candido, James Maningo, et al. 2023 · 2023
Learning to make rare and complex diagnoses with generative ai assistance: Qualitative study of popular large language models
Tassallah Abdullahi, Ritambhara Singh, Carsten Eickhoff, et al. 2024 · 2024
Closest in time.
Introducing claude
Anthropic. 2023 · 2024
Closest in time.
Benchmarking large language models on answering and explaining challenging medical questions
Hanjie Chen, Zhouxiang Fang, Yash Singla, and Mark Dredze. 2024 · 2024
Closest in time.
Medalign: A clinician-generated dataset for instruction following with electronic medical records
Scott L. Fleming, Alejandro Lozano, William J. Haberkorn, Jenelle A. Jindal, Eduardo Reis, Rahul Thapa, Louis Blankemeier, Julian Z. Genkins, Ethan Steinberg, Ashwin Nayak, Birju Patel, Chia-Chun Chiang, Alison Callahan, Zepeng Huo, Sergios Gatidis, Scott Adams, Oluseyi Fayanju, Shreya J. Shah, Thomas Savage, Ethan Goh, Akshay S. Chaudhari, Nima Aghaeepour, Christopher Sharp, Michael A. Pfeffer, Percy Liang, Jonathan H. Chen, Keith E. Morse, Emma P. Brunskill, Jason A. Fries, and Nigam H. Shah. 2024 · 2024
Closest in time.
MedQA-SWE - a clinical question & answer dataset for Swedish
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Benchmarking large language models on cmexam - a comprehensive chinese medical exam dataset
Junling Liu, Peilin Zhou, Yining Hua, Dading Chong, Zhongyu Tian, Andrew Liu, Helin Wang, Chenyu You, Zhenhua Guo, LEI ZHU, and Michael Lingzhi Li. 2023 · 2023
Cited alongside, same era.
Afrispeech-200: Pan-african accented speech dataset for clinical and general domain asr
Tobi Olatunji, Tejumade Afonja, Aditya Yadavalli, Chris Chinenye Emezue, Sahib Singh, Bonaventure FP Dossou, Joanne Osuchukwu, Salomey Osei, Atnafu Lambebo Tonja, Naome Etori, et al. 2023 · 2023
Cited alongside, same era.
Evaluating large language models for use in healthcare: A framework for translational value assessment
Sandeep Reddy. 2023 · 2023
Cited alongside, same era.
Towards expert-level medical question answering with large language models
Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Le Hou, Kevin Clark, Stephen Pfohl, Heather Cole-Lewis, Darlene Neal, et al. 2023 · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. 2023 · 2023
Cited alongside, same era.
Med-halt: Medical domain hallucination test for large language models
Logesh Kumar Umapathi, Ankit Pal, and Malaikannan Sankarasubbu. 2023 · 2023
Cited alongside, same era.
Large language models in health care: Development, applications, and challenges
Rui Yang, Ting Fang Tan, Wei Lu, Arun James Thirunavukarasu, Daniel Shu Wei Ting, and Nan Liu. 2023 · 2023
Cited alongside, same era.
Niclas Hertzberg and Anna Lokrantz. 2024 · 2024
Closest in time.
A comprehensive evaluation of large language models on benchmark biomedical text processing tasks
Israt Jahan, Md Tahmid Rahman Laskar, Chun Peng, and Jimmy Xiangji Huang. 2024 · 2024
Closest in time.
openlifescienceai/open_medical_llm_leaderboard
Ankit Pal, Pasquale Minervini, Andreas Geert Motzfeldt, Aryo Pradipta Gema, and Beatrice Alex. 2024 · 2024
Closest in time.
Openbiollms: Advancing open-source large language models for healthcare and life sciences
Ankit Pal and Malaikannan Sankarasubbu. 2024 · 2024
Closest in time.
A toolbox for surfacing health equity harms and biases in large language models
Stephen R Pfohl, Heather Cole-Lewis, Rory Sayres, Darlene Neal, Mercy Asiedu, Awa Dieng, Nenad Tomasev, Qazi Mamunur Rashid, Shekoofeh Azizi, Negar Rostamzadeh, et al. 2024 · 2024
Closest in time.
Climategpt: Towards ai synthesizing interdisciplinary research on climate change
David Thulke, Yingbo Gao, Petrus Pelser, Rein Brune, Rricha Jalota, Floris Fok, Michael Ramos, Ian van Wyk, Abdallah Nasir, Hayden Goldstein, et al. 2024 · 2024
Closest in time.
Towards conversational diagnostic ai
Tao Tu, Anil Palepu, Mike Schaekermann, Khaled Saab, Jan Freyberg, Ryutaro Tanno, Amy Wang, Brenna Li, Mohamed Amin, Nenad Tomasev, et al. 2024 · 2024
Closest in time.
Rouge-sem: Better evaluation of summarization using rouge combined with semantics
Ming Zhang, Chengzhang Li, Meilin Wan, Xuejun Zhang, and Qingwei Zhao. 2024 · 2024
Closest in time.
Large language models are not robust multiple choice selectors
Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. 2024 · 2024
Closest in time.