Fetching the paper…
Reading the bibliography…
In recent years, machine learning models have rapidly become better at generating clinical consultation notes; yet, there is little work on how to properly evaluate the generated consultation notes to understand the impact they may have on both the clinician using them and the patient's clinical safety.
Binary codes capable of correcting deletions, insertions, and reversals
Vladimir I Levenshtein et al. 1966 · 1966
Earlier work this paper cites.
A new quantitative quality measure for machine translation systems
Keh-Yih Su, Ming-Wen Wu, and Jing-Shin Chang. 1992 · 1992
Earlier work this paper cites.
Snomed rt: a reference terminology for health care
Kent A Spackman, Keith E Campbell, and Roger A Côté. 1997 · 1997
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
From wer and ril to mer and wil: improved evaluation measures for connected speech recognition
Andrew C. Morris, Viktoria Maier, and Phil D. Green. 2004 · 2004
Earlier work this paper cites.
Evaluation of an nlg system using post-edit data: Lessons learnt
Somayajulu Sripada, Ehud Reiter, and Lezan Hawizy. 2005 · 2005
Earlier work this paper cites.
Spearman rank correlation
Jerrold H Zar. 2005 · 2005
Earlier work this paper cites.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Earlier work this paper cites.
Statistics (international student edition)
David Freedman, Robert Pisani, and Roger Purves. 2007 · 2007
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with high levels of correlation with human judgments
Alon Lavie and Abhaya Agarwal. 2007 · 2007
Earlier work this paper cites.
A survey of evaluation metrics used for nlg systems
Ananya B Sai, Akash Kumar Mohankumar, and Mitesh M Khapra. 2020 · 2008
Earlier work this paper cites.
Get another label? improving data quality and data mining using multiple, noisy labelers
Victor S Sheng, Foster Provost, and Panagiotis G Ipeirotis. 2008 · 2008
Earlier work this paper cites.
Bootstrapping dialog systems with word embeddings
Gabriel Forgues, Joelle Pineau, Jean-Marie Larchevêque, and Réal Tremblay. 2014 · 2014
Cited alongside, same era.
From word embeddings to document distances
Matt Kusner, Yu Sun, Nicholas Kolkin, and Kilian Weinberger. 2015 · 2015
Cited alongside, same era.
chrf: character n-gram f-score for automatic mt evaluation
Maja Popović. 2015 · 2015
Cited alongside, same era.
Abstractive text summarization using sequence-to-sequence rnns and beyond
Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Çağlar Gulçehre, and Bing Xiang. 2016 · 2016
Cited alongside, same era.
The essential soap note in an ehr age
Patricia F Pearce, Laurie Anne Ferguson, Gwen S George, and Cynthia A Langford. 2016 · 2016
Cited alongside, same era.
Tethered to the ehr: primary care physician workload assessment using ehr event log data and time-motion observations
Brian G Arndt, John W Beasley, Michelle D Watkinson, Jonathan L Temte, Wen-Jan Tuan, Christine A Sinsky, and Valerie J Gilchrist. 2017 · 2017
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 2019
Later among the works it cites.
Moverscore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M Meyer, and Steffen Eger. 2019 · 2019
Later among the works it cites.
Generating medical reports from patient-doctor conversations using sequence-to-sequence models
Seppo Enarvi, Marilisa Amoia, Miguel Del-Agua Teba, Brian Delaney, Frank Diehl, Stefan Hahn, Kristina Harris, Liam McGrath, Yue Pan, Joel Pinto, et al. 2020 · 2020
Later among the works it cites.
Dr. summarize: Global summarization of medical dialogue by exploiting local structures
Anirudh Joshi, Namit Katariya, Xavier Amatriain, and Anitha Kannan. 2020 · 2020
Later among the works it cites.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Shikhar Sharma, Layla El Asri, Hannes Schulz, and Jeremie Zumer. 2017 · 2017
Cited alongside, same era.
Extractive summarization of ehr discharge notes
Emily Alsentzer and Anne Kim. 2018 · 2018
Cited alongside, same era.
Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St John, Noah Constant, Mario Guajardo-Céspedes, Steve Yuan, Chris Tar, et al. 2018 · 2018
Cited alongside, same era.
Content analysis: An introduction to its methodology
Klaus Krippendorff. 2018 · 2018
Cited alongside, same era.
Learning to summarize radiology findings
Yuhao Zhang, Daisy Yi Ding, Tianpei Qian, Christopher D Manning, and Curtis P Langlotz. 2018 · 2018
Cited alongside, same era.
Topic-aware pointer-generator networks for summarizing spoken conversations
Zhengyuan Liu, Angela Ng, Sheldon Lee, Ai Ti Aw, and Nancy F Chen. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Optimizing the factual correctness of a summary: A study of summarizing radiology reports
Yuhao Zhang, Derek Merck, Emily Tsai, Christopher D Manning, and Curtis Langlotz. 2020 · 2020
Later among the works it cites.
Medically aware gpt-3 as a data generator for medical dialogue summarization
Bharath Chintagunta, Namit Katariya, Xavier Amatriain, and Anitha Kannan. 2021 · 2021
Later among the works it cites.
Summeval: Re-evaluating summarization evaluation
Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Later among the works it cites.
Generating SOAP notes from doctor-patient conversations using modular summarization techniques
Kundan Krishna, Sopan Khosla, Jeffrey Bigham, and Zachary C. Lipton. 2021 · 2021
Later among the works it cites.
A preliminary study on evaluating consultation notes with post-editing
Francesco Moramarco, Alex Papadopoulos Korfiatis, Aleksandar Savkov, and Ehud Reiter. 2021 · 2021
Later among the works it cites.
A reproduction study of an annotation-based human evaluation of mt outputs
Maja Popović and Anya Belz. 2021 · 2021
Later among the works it cites.
Towards automating medical scribing: Clinic visit dialogue2note sentence alignment and snippet summarization
Wen-wai Yim and Meliha Yetisgen-Yildiz. 2021 · 2021
Later among the works it cites.
(in press): Primock57: A dataset of primary care mock consultations
Alex Papadopoulos Korfiatis, Francesco Moramarco, Radmila Sarac, and Aleksandar Savkov. 2022 · 2022
Closest in time.