Fetching the paper…
Reading the bibliography…
While human evaluation remains best practice for accurately judging the faithfulness of automatically-generated summaries, few solutions exist to address the increased difficulty and workload when evaluating long-form summaries.
Measuring nominal scale agreement among many raters
Joseph L Fleiss. 1971 · 1971
Earlier work this paper cites.
An introduction to the bootstrap
Robert J Tibshirani and Bradley Efron. 1993 · 1993
Earlier work this paper cites.
Okapi at trec-3
Stephen E Robertson, Steve Walker, Susan Jones, Micheline M Hancock-Beaulieu, Mike Gatford, et al. 1995 · 1995
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Evaluating content selection in summarization: The pyramid method
Ani Nenkova and Rebecca Passonneau. 2004 · 2004
Earlier work this paper cites.
Free-marginal multirater kappa (multirater k [free]): An alternative to fleiss’ fixed-marginal multirater kappa
Justus J Randolph. 2005 · 2005
Earlier work this paper cites.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Earlier work this paper cites.
Non-expert evaluation of summarization systems is risky
Dan Gillick and Yang Liu. 2010 · 2010
Earlier work this paper cites.
Inequalities between multi-rater kappas
Matthijs J Warrens. 2010 · 2010
Earlier work this paper cites.
Fact checking: Task definition and dataset construction
Andreas Vlachos and Sebastian Riedel. 2014 · 2014
Earlier work this paper cites.
Results of the WMT16 metrics shared task
Ondřej Bojar, Yvette Graham, Amir Kamran, and Miloš Stanojević. 2016 · 2016
Earlier work this paper cites.
Copyright and fair use
OGC of Harvard. 2016 · 2016
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence RNNs and beyond
Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Caglar Gulcehre, and Bing Xiang. 2016 · 2016
Earlier work this paper cites.
A discourse-aware attention model for abstractive summarization of long documents
Arman Cohan, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Seokhwan Kim, Walter Chang, and Nazli Goharian. 2018 · 2018
Earlier work this paper cites.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018 · 2018
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018 · 2018
Earlier work this paper cites.
Multi-news: A large-scale multi-document summarization dataset and abstractive hierarchical model
Alexander Fabbri, Irene Li, Tianwei She, Suyi Li, and Dragomir Radev. 2019 · 2019
Earlier work this paper cites.
Ranking generated summaries by correctness: An interesting but challenging application for natural language inference
Tobias Falke, Leonardo F. R. Ribeiro, Prasetya Ajie Utama, Ido Dagan, and Iryna Gurevych. 2019 · 2019
Earlier work this paper cites.
HighRES: Highlight-based reference-less evaluation of summarization
Hardy Hardy, Shashi Narayan, and Andreas Vlachos. 2019 · 2019
Earlier work this paper cites.
BillSum: A corpus for automatic summarization of US legislation
Anastassia Kornilova and Vladimir Eidelman. 2019 · 2019
Earlier work this paper cites.
Neural text summarization: A critical evaluation
Wojciech Kryscinski, Nitish Shirish Keskar, Bryan McCann, Caiming Xiong, and Richard Socher. 2019 · 2019
Earlier work this paper cites.
Crowdsourcing lightweight pyramids for manual summary evaluation
Ori Shapira, David Gabay, Yang Gao, Hadar Ronen, Ramakanth Pasunuru, Mohit Bansal, Yael Amsterdamer, and Ido Dagan. 2019 · 2019
Cited alongside, same era.
Beyond BLEU:training neural machine translation with semantic similarity
John Wieting, Taylor Berg-Kirkpatrick, Kevin Gimpel, and Graham Neubig. 2019 · 2019
Cited alongside, same era.
STORIUM: A Dataset and Evaluation Platform for Machine-in-the-Loop Story Generation
Nader Akoury, Shufan Wang, Josh Whiting, Stephen Hood, Nanyun Peng, and Mohit Iyyer. 2020 · 2020
Cited alongside, same era.
Re-evaluating evaluation in text summarization
Manik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu, and Graham Neubig. 2020 · 2020
Cited alongside, same era.
SacreROUGE: An Open-Source Library for Using and Developing Summarization Evaluation Metrics
Daniel Deutsch and Dan Roth. 2020 · 2020
Cited alongside, same era.
Leveraging information bottleneck for scientific document summarization
Jiaxin Ju, Ming Liu, Huan Yee Koh, Yuan Jin, Lan Du, and Shirui Pan. 2021 · 2021
Later among the works it cites.
The perils of using Mechanical Turk to evaluate open-ended text generation
Marzena Karpinska, Nader Akoury, and Mohit Iyyer. 2021 · 2021
Later among the works it cites.
Hurdles to progress in long-form question answering
Kalpesh Krishna, Aurko Roy, and Mohit Iyyer. 2021 · 2021
Later among the works it cites.
Booksum: A collection of datasets for long-form narrative summarization
Wojciech Kryściński, Nazneen Rajani, Divyansh Agarwal, Caiming Xiong, and Dragomir Radev. 2021 · 2021
Later among the works it cites.
Automated fact-checking for assisting human fact-checkers
Preslav Nakov, David Corney, Maram Hasanain, Firoj Alam, Tamer Elsayed, Alberto Barrón-Cedeño, Paolo Papotti, Shaden Shaar, and Giovanni Da San Martino. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Angela Fan, Aleksandra Piktus, Fabio Petroni, Guillaume Wenzek, Marzieh Saeidi, Andreas Vlachos, Antoine Bordes, and Sebastian Riedel. 2020 · 2020
Cited alongside, same era.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2020
Cited alongside, same era.
Evaluating the factual consistency of abstractive text summarization
Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020 · 2020
Cited alongside, same era.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Cited alongside, same era.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Cited alongside, same era.
Copyright and fair use guidelines
Libraries at UMGC. 2020 · 2020
Cited alongside, same era.
Fact or fiction: Verifying scientific claims
David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
Measuring attribution in natural language generation models
Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Michael Collins, Dipanjan Das, Slav Petrov, Gaurav Singh Tomar, Iulia Turc, and David Reitter. 2021 · 2021
Later among the works it cites.
SummVis: Interactive visual analysis of models, data, and evaluation for text summarization
Jesse Vig, Wojciech Kryscinski, Karan Goel, and Nazneen Rajani. 2021 · 2021
Later among the works it cites.
The statistical advantage of automatic NLG metrics at the system level
Johnny Wei and Robin Jia. 2021 · 2021
Later among the works it cites.
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Later among the works it cites.
Finding a balanced degree of automation for summary evaluation
Shiyue Zhang and Mohit Bansal. 2021 · 2021
Later among the works it cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2020 · 2021
Later among the works it cites.
SummScreen: A dataset for abstractive screenplay summarization
Mingda Chen, Zewei Chu, Sam Wiseman, and Kevin Gimpel. 2022 · 2022
Later among the works it cites.
QAFactEval: Improved QA-based factual consistency evaluation for summarization
Alexander Fabbri, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2022 · 2022
Later among the works it cites.
Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text
Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. 2022 · 2022
Later among the works it cites.
Snac: Coherence error detection for narrative summarization
Tanya Goyal, Junyi Jessy Li, and Greg Durrett. 2022 · 2022
Later among the works it cites.
LongT5: Efficient text-to-text transformer for long sequences
Mandy Guo, Joshua Ainslie, David Uthus, Santiago Ontanon, Jianmo Ni, Yun-Hsuan Sung, and Yinfei Yang. 2022 · 2022
Later among the works it cites.
SummaC: Re-visiting NLI-based models for inconsistency detection in summarization
Philippe Laban, Tobias Schnabel, Paul N. Bennett, and Marti A. Hearst. 2022 · 2022
Later among the works it cites.
Revisiting the gold standard: Grounding summarization evaluation with robust human evaluation
Yixin Liu, Alexander R Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, et al. 2022 · 2022
Later among the works it cites.
Investigating crowdsourcing protocols for evaluating the factual consistency of summaries
Xiangru Tang, Alexander Fabbri, Haoran Li, Ziming Mao, Griffin Adams, Borui Wang, Asli Celikyilmaz, Yashar Mehdad, and Dragomir Radev. 2022 · 2022
Later among the works it cites.
MultiVerS: Improving scientific claim verification with weak supervision and full-document context
David Wadden, Kyle Lo, Lucy Wang, Arman Cohan, Iz Beltagy, and Hannaneh Hajishirzi. 2022 · 2022
Later among the works it cites.
Squality: Building a long-document summarization dataset the hard way
Alex Wang, Richard Yuanzhe Pang, Angelica Chen, Jason Phang, and Samuel R Bowman. 2022 · 2022
Later among the works it cites.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2022 · 2022
Later among the works it cites.