Fetching the paper…
Reading the bibliography…
This paper describes our approach to the MEDIQA-CORR shared task, which involves error detection and correction in clinical notes curated by medical professionals.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Earlier work this paper cites.
SemEval-2020 task 4: Commonsense validation and explanation
Cunxiang Wang, Shuailong Liang, Yili Jin, Yilong Wang, Xiaodan Zhu, and Yue Zhang. 2020 · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021 · 2021
Earlier work this paper cites.
An investigation of evaluation metrics for automated medical note generation
Asma Ben Abacha, Wen wai Yim, George Michalopoulos, and Thomas Lin. 2023 · 2023
Earlier work this paper cites.
The future landscape of large language models in medicine
Jan Clusmann, Fiona R Kolbinger, et al. 2023 · 2023
Cited alongside, same era.
Kai He, Rui Mao, Qika Lin, Yucheng Ruan, Xiang Lan, Mengling Feng, and Erik Cambria. 2023 · 2023
Cited alongside, same era.
Natural language processing in electronic health records in relation to healthcare decision-making: A systematic review
Elias Hossain, Rajib Rana, Niall Higgins, Jeffrey Soar, Prabal Datta Barua, Anthony R. Pisani, and Kathryn Turner. 2023 · 2023
Cited alongside, same era.
Large language models in healthcare and medical domain: A review
Zabir Al Nazi and Wei Peng. 2023 · 2023
Cited alongside, same era.
Can generalist foundation models outcompete special-purpose tuning? case study in medicine
A comprehensive capability analysis of gpt-3 and gpt-3.5 series models
Junjie Ye, Xuanting Chen, Nuo Xu, Can Zu, Zekai Shao, Shichun Liu, Yuhan Cui, Zeyang Zhou, Chao Gong, Yang Shen, Jie Zhou, Siming Chen, Tao Gui, Qi Zhang, and Xuanjing Huang. 2023 · 2023
Later among the works it cites.
A survey of large language models in medicine: Progress, application, and challenge
Hongjian Zhou, Fenglin Liu, Boyang Gu, Xinyu Zou, Jinfa Huang, Jinge Wu, Yiru Li, Sam S. Chen, Peilin Zhou, Yining Hua Junling Liu, Chengfeng Mao, Xian Wu, Yefeng Zheng, Lei Clifton, Zheng Li, Jiebo Luo, and David A. Clifton. 2023 · 2023
Later among the works it cites.
Medec: A benchmark for medical error detection and correction in clinical notes
Asma Ben Abacha, Wen wai Yim, Velvin Fu, Zhaoyi Sun, Meliha Yetisgen, Fei Xia, and Thomas Lin. 2024 · 2024
Closest in time.
The claude 3 model family: Opus, sonnet, haiku
Anthropic and others. 2024 · 2024
Closest in time.
Overview of the mediqa-corr 2024 shared task on medical error detection and correction
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Harsha Nori, Yin Tat Lee, Sheng Zhang, Dean Carignan, Richard Edgar, Nicolo Fusi, Nicholas King, Jonathan Larson, Yuanzhi Li, Weishung Liu, Renqian Luo, Scott Mayer McKinney, Robert Osazuwa Ness, Hoifung Poon, Tao Qin, Naoto Usuyama, Chris White, and Eric Horvitz. 2023 · 2023
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023 · 2023
Cited alongside, same era.
Asma Ben Abacha, Wen wai Yim, Velvin Fu, Zhaoyi Sun, Fei Xia, and Meliha Yetisgen. 2024 · 2024
Closest in time.
OpenAI, Josh Achiam, et al. 2024 · 2024
Closest in time.