Fetching the paper…
Reading the bibliography…
While there has been significant development of models for Plain Language Summarization (PLS), evaluation remains a challenge.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
R Thomas McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 1902
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 1904
Earlier work this paper cites.
A survey of automated methods for biomedical text simplification
Brian Ondov, Kush Attal, and Dina Demner-Fushman. 2022 · 1988
Earlier work this paper cites.
Revising social studies text from a text-processing perspective: Evidence of improved comprehensibility
Isabel L Beck, Margaret G McKeown, Gale M Sinatra, and Jane A Loxterman. 1991 · 1991
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2005
Earlier work this paper cites.
Evidence-based patient choice: a prostate cancer decision aid in plain language
Margaret Holmes-Rovner, Sue Stableford, Angela Fagerlin, John T Wei, Rodney L Dunn, Janet Ohene-Frempong, Karen Kelly-Blake, and David R Rovner. 2005 · 2005
Earlier work this paper cites.
Inter-coder agreement for computational linguistics
Ron Artstein and Massimo Poesio. 2008 · 2008
Earlier work this paper cites.
The acl anthology reference corpus: A reference dataset for bibliographic research in computational linguistics
Steven Bird, Robert Dale, Bonnie J Dorr, Bryan R Gibson, Mark Thomas Joseph, Min-Yen Kan, Dongwon Lee, Brett Powley, Dragomir R Radev, Yee Fan Tan, et al. 2008 · 2008
Earlier work this paper cites.
Text simplification tools: Using machine learning to discover features that identify difficult text
David Kauchak, Obay Mouradi, Christopher Pentoney, and Gondy Leroy. 2014 · 2014
Earlier work this paper cites.
Lay summaries needed to enhance science communication
Lauren M Kuehne and Julian D Olden. 2015 · 2015
Earlier work this paper cites.
Optimizing statistical machine translation for text simplification
Wei Xu, Courtney Napoles, Ellie Pavlick, Quanze Chen, and Chris Callison-Burch. 2016 · 2016
Earlier work this paper cites.
The role of surface, semantic and grammatical features on simplification of spanish medical texts: A user study
Partha Mukherjee, Gondy Leroy, David Kauchak, Brianda Armenta Navarrete, Damian Y Diaz, and Sonia Colina. 2017 · 2017
Earlier work this paper cites.
Next-generation metrics for monitoring genetic erosion within populations of conservation concern
Gregoire Leroy, Emma L Carroll, Mike W Bruford, J Andrew DeWoody, Allan Strand, Lisette Waits, and Jinliang Wang. 2018 · 2018
Earlier work this paper cites.
Bleu is not suitable for the evaluation of text simplification
Elior Sulem, Omri Abend, and Ari Rappoport. 2018 · 2018
Earlier work this paper cites.
EASSE: Easier automatic sentence simplification evaluation
Fernando Alva-Manchego, Louis Martin, Carolina Scarton, and Lucia Specia. 2019 · 2019
Earlier work this paper cites.
Highres: Highlight-based reference-less evaluation of summarization
Hardy Hardy, Shashi Narayan, and Andreas Vlachos. 2019 · 2019
Earlier work this paper cites.
Domain agnostic real-valued specificity prediction
Wei-Jen Ko, Greg Durrett, and Junyi Jessy Li. 2019 · 2019
Earlier work this paper cites.
Are red roses red? evaluating consistency of question-answering models
Marco Tulio Ribeiro, Carlos Guestrin, and Sameer Singh. 2019 · 2019
Cited alongside, same era.
Leveraging social media for medical text simplification
Nikhil Pattisapu, Nishant Prabhu, Smriti Bhati, and Vasudeva Varma. 2020 · 2020
Cited alongside, same era.
Assessing the benchmarking capacity of machine reading comprehension datasets
Saku Sugawara, Pontus Stenetorp, Kentaro Inui, and Akiko Aizawa. 2020 · 2020
Cited alongside, same era.
All that’s ‘human’ is not gold: Evaluating human evaluation of generated text
Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A. Smith. 2021 · 2021
Cited alongside, same era.
Towards question-answering as an automatic metric for evaluating the content quality of a summary
Daniel Deutsch, Tania Bedrax-Weiss, and Dan Roth. 2021 · 2021
Cited alongside, same era.
Template and guidance for writing a cochrane plain language summary
Nicole Pitcher, Denise Mitchell, and Carolyn Hughes. 2022 · 2022
Later among the works it cites.
A survey of evaluation metrics used for nlg systems
Ananya B Sai, Akash Kumar Mohankumar, and Mitesh M Khapra. 2022 · 2022
Later among the works it cites.
Plain language summaries: A systematic review of theory, guidelines and empirical research
Marlene Stoll, Martin Kerwer, Klaus Lieb, and Anita Chasiotis. 2022 · 2022
Later among the works it cites.
Generating scientific claims for zero-shot scientific fact checking
Dustin Wright, David Wadden, Kyle Lo, Bailey Kuehl, Arman Cohan, Isabelle Augenstein, and Lucy Lu Wang. 2022 · 2022
Later among the works it cites.
Improving meta-learning for low-resource text classification and generation via memory imitation
Yingxiu Zhao, Zhiliang Tian, Huaxiu Yao, Yinhe Zheng, Dongkyu Lee, Yiping Song, Jian Sun, and Nevin Zhang. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ashwin Devaraj, Iain Marshall, Byron C Wallace, and Junyi Jessy Li. 2021 · 2021
Cited alongside, same era.
Summeval: Re-evaluating summarization evaluation
Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Cited alongside, same era.
Go figure: A meta evaluation of factuality in summarization
Saadia Gabriel, Asli Celikyilmaz, Rahul Jha, Yejin Choi, and Jianfeng Gao. 2021 · 2021
Cited alongside, same era.
Automated lay language summarization of biomedical scientific reviews
Yue Guo, Wei Qiu, Yizhong Wang, and Trevor Cohen. 2021 · 2021
Cited alongside, same era.
Summarization of legal documents: Where are we now and the way forward
Deepali Jain, Malaya Dutta Borah, and Anupam Biswas. 2021 · 2021
Cited alongside, same era.
Understanding factuality in abstractive summarization with frank: A benchmark for factuality metrics
Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021 · 2021
Cited alongside, same era.
Perturbation checklists for evaluating nlg evaluation metrics
Ananya B Sai, Tanay Dixit, Dev Yashpal Sheth, Sreyas Mohan, and Mitesh M Khapra. 2021 · 2021
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Closest in time.
A dataset for plain language adaptation of biomedical abstracts
Kush Attal, Brian Ondov, and Dina Demner-Fushman. 2023 · 2023
Closest in time.
Paper plain: Making medical research papers approachable to healthcare consumers with natural language processing
Tal August, Lucy Lu Wang, Jonathan Bragg, Marti A Hearst, Andrew Head, and Kyle Lo. 2023 · 2023
Closest in time.
Menli: Robust evaluation metrics from natural language inference
Yanran Chen and Steffen Eger. 2023 · 2023
Closest in time.
Human-like summarization evaluation with chatgpt
Mingqi Gao, Jie Ruan, Renliang Sun, Xunjian Yin, Shiping Yang, and Xiaojun Wan. 2023 · 2023
Closest in time.
Domain-driven and discourse-guided scientific summarisation
Tomas Goldsack, Zhihao Zhang, Chenghua Lin, and Carolina Scarton. 2023 · 2023
Closest in time.
On the blind spots of model-based evaluation metrics for text generation
Tianxing He, Jingyu Zhang, Tianle Wang, Sachin Kumar, Kyunghyun Cho, James Glass, and Yulia Tsvetkov. 2023 · 2023
Closest in time.
Longeval: Guidelines for human evaluation of faithfulness in long-form summarization
Kalpesh Krishna, Erin Bransom, Bailey Kuehl, Mohit Iyyer, Pradeep Dasigi, Arman Cohan, and Kyle Lo. 2023 · 2023
Closest in time.
Napss: Paragraph-level medical text simplification via narrative prompting and sentence-matching summarization
Junru Lu, Jiazheng Li, Byron C. Wallace, Yulan He, and Gabriele Pergola. 2023 · 2023
Closest in time.
Chatgpt as a factual inconsistency evaluator for abstractive text summarization
Zheheng Luo, Qianqian Xie, and Sophia Ananiadou. 2023 · 2023
Closest in time.
Lens: A learnable evaluation metric for text simplification
Mounica Maddela, Yao Dou, David Heineman, and Wei Xu. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Closest in time.
Automated metrics for medical multi-document summarization disagree with human evaluations
Lucy Lu Wang, Yulia Otmakhova, Jay DeYoung, Thinh Hung Truong, Bailey Kuehl, Erin Bransom, and Byron Wallace. 2023 · 2023
Closest in time.
Retrieval augmentation of large language models for lay language generation
Yue Guo, Wei Qiu, Gondy Leroy, Sheng Wang, and Trevor Cohen. 2024 · 2024
Closest in time.
Measuring text difficulty using parse-tree frequency
David Kauchak, Gondy Leroy, and Alan Hogue. 2017 · 2088
Closest in time.