Fetching the paper…
Reading the bibliography…
A desirable property of a reference-based evaluation metric that measures the content quality of a summary is that it should estimate how much information that summary has in common with a reference.
ROUGE: A Package for Automatic Evaluation of Summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Evaluating Content Selection in Summarization: The Pyramid Method
Ani Nenkova and Rebecca J. Passonneau. 2004 · 2004
Earlier work this paper cites.
Automated Summarization Evaluation with Basic Elements
Eduard H. Hovy, Chin-Yew Lin, Liang Zhou, and Junichi Fukumoto. 2006 · 2006
Earlier work this paper cites.
SummEval: Re-evaluating Summarization Evaluation
Alexander R. Fabbri, Wojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir R. Radev. 2020 · 2007
Earlier work this paper cites.
Overview of the TAC 2008 Update Summarization Task
Hoa Trang Dang and Karolina Owczarzak. 2008 · 2008
Earlier work this paper cites.
Summarization System Evaluation Revisited: N-Gram Graphs
George Giannakopoulos, Vangelis Karkaletsis, George A. Vouros, and Panagiotis Stamatopoulos. 2008 · 2008
Earlier work this paper cites.
Summarization Evaluation Using Transformed Basic Elements
Stephen Tratz and Eduard H. Hovy. 2008 · 2008
Earlier work this paper cites.
Overview of the TAC 2009 Summarization Track
Hoa Trang Dang and Karolina Owczarzak. 2009 · 2009
Earlier work this paper cites.
Daniel Deutsch and Dan Roth. 2020 · 2010
Earlier work this paper cites.
Automatically Assessing Machine Summary Content Without a Gold Standard
Annie Louis and Ani Nenkova. 2013 · 2013
Earlier work this paper cites.
Meteor Universal: Language Specific Translation Evaluation for Any Target Language
Michael J. Denkowski and Alon Lavie. 2014 · 2014
Cited alongside, same era.
Abstractive Text Summarization using Sequence-to-sequence RNNs and Beyond
Ramesh Nallapati, Bowen Zhou, Cícero Nogueira dos Santos, Çaglar Gülçehre, and Bing Xiang. 2016 · 2016
Cited alongside, same era.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
PEAK: Pyramid Evaluation via Automated Knowledge Extraction
Qian Yang, Rebecca J. Passonneau, and Gerard de Melo. 2016 · 2016
Cited alongside, same era.
Learning to Score System Summaries for Better Content Selection Evaluation
Maxime Peyrard, Teresa Botschen, and Iryna Gurevych. 2017 · 2017
Cited alongside, same era.
Transforming Question Answering Datasets Into Natural Language Inference Datasets
Crowdsourcing Lightweight Pyramids for Manual Summary Evaluation
Ori Shapira, David Gabay, Yang Gao, Hadar Ronen, Ramakanth Pasunuru, Mohit Bansal, Yael Amsterdamer, and Ido Dagan. 2019 · 2019
Later among the works it cites.
MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019 · 2019
Later among the works it cites.
Re-evaluating Evaluation in Text Summarization
Manik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu, and Graham Neubig. 2020 · 2020
Closest in time.
MOCHA: A Dataset for Training and Evaluating Generative Reading Comprehension Metrics
Anthony Chen, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2020 · 2020
Closest in time.
ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020 · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dorottya Demszky, Kelvin Guu, and Percy Liang. 2018 · 2018
Cited alongside, same era.
Ranking Sentences for Extractive Summarization with Reinforcement Learning
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018 · 2018
Cited alongside, same era.
Know What You Don’t Know: Unanswerable Questions for SQuAD
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Cited alongside, same era.
Question Answering as an Automatic Evaluation Metric for News Article Summarization
Matan Eyal, Tal Baumel, and Michael Elhadad. 2019 · 2019
Cited alongside, same era.
Automated Pyramid Summarization Evaluation
Yanjun Gao, Chen Sun, and Rebecca J. Passonneau. 2019 · 2019
Cited alongside, same era.
FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive Summarization
Esin Durmus, He He, and Mona Diab. 2020 · 2020
Closest in time.
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Closest in time.
Asking and Answering Questions to Evaluate the Factual Consistency of Summaries
Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020 · 2020
Closest in time.
BERTScore: Evaluating Text Generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Closest in time.
A Statistical Analysis of Summarization Evaluation Metrics using Resampling Methods
Daniel Deutsch, Rotem Dror, and Dan Roth. 2021 · 2021
Closest in time.