Fetching the paper…
Reading the bibliography…
Current pre-trained models applied to summarization are prone to factual inconsistencies which either misrepresent the source text or introduce extraneous information.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020 · 1919
Earlier work this paper cites.
Best-worst scaling: A model for the largest difference judgments
Jordan J Louviere and George G Woodworth. 1991 · 1991
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Non-expert evaluation of summarization systems is risky
Dan Gillick and Yang Liu. 2010 · 2010
Earlier work this paper cites.
Ffci: A framework for interpretable automatic evaluation of summarization
Fajri Koto, Jey Han Lau, and Timothy Baldwin. 2020 · 2011
Earlier work this paper cites.
Computing krippendorff’s alpha-reliability
Klaus Krippendorff. 2011 · 2011
Earlier work this paper cites.
Crowdsourcing research opportunities: lessons from natural language processing
Marta Sabou, Kalina Bontcheva, and Arno Scharl. 2012 · 2012
Earlier work this paper cites.
Analyzing the capabilities of crowdsourcing services for text summarization
Elena Lloret, Laura Plaza, and Ahmet Aker. 2013 · 2013
Earlier work this paper cites.
Automatically assessing machine summary content without a gold standard
Annie Louis and Ani Nenkova. 2013 · 2013
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomás Kociský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence RNNs and beyond
Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Caglar Gulcehre, and Bing Xiang. 2016 · 2016
Earlier work this paper cites.
Best-worst scaling more reliable than rating scales: A case study on sentiment intensity annotation
Svetlana Kiritchenko and Saif Mohammad. 2017 · 2017
Earlier work this paper cites.
Affective neural response generation
Nabiha Asghar, Pascal Poupart, Jesse Hoey, Xin Jiang, and Lili Mou. 2018 · 2018
Earlier work this paper cites.
Scoring best-worst data in unbalanced many-item designs, with applications to crowdsourcing semantic judgments
Geoff Hollis. 2018 · 2018
Cited alongside, same era.
When is best-worst best? a comparison of best-worst scaling, numeric estimation, and rating scales for collection of semantic norms
Geoff Hollis and Chris Westbury. 2018 · 2018
Cited alongside, same era.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018 · 2018
Cited alongside, same era.
Unified language model pre-training for natural language understanding and generation
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019 · 2019
Cited alongside, same era.
Pre-trained language model representations for language generation
Sergey Edunov, Alexei Baevski, and Michael Auli. 2019 · 2019
Cited alongside, same era.
Factual error correction for abstractive summarization models
Meng Cao, Yue Dong, Jiapeng Wu, and Jackie Chi Kit Cheung. 2020 · 2020
Later among the works it cites.
FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization
Esin Durmus, He He, and Mona Diab. 2020 · 2020
Later among the works it cites.
Evaluating the factual consistency of abstractive text summarization
Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020 · 2020
Later among the works it cites.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
ProphetNet: Predicting future n-gram for sequence-to-SequencePre-training
Weizhen Qi, Yu Yan, Yeyun Gong, Dayiheng Liu, Nan Duan, Jiusheng Chen, Ruofei Zhang, and Ming Zhou. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Question answering as an automatic evaluation metric for news article summarization
Matan Eyal, Tal Baumel, and Michael Elhadad. 2019 · 2019
Cited alongside, same era.
Ranking generated summaries by correctness: An interesting but challenging application for natural language inference
Tobias Falke, Leonardo F. R. Ribeiro, Prasetya Ajie Utama, Ido Dagan, and Iryna Gurevych. 2019 · 2019
Cited alongside, same era.
HighRES: Highlight-based reference-less evaluation of summarization
Hardy Hardy, Shashi Narayan, and Andreas Vlachos. 2019 · 2019
Cited alongside, same era.
Text summarization with pretrained encoders
Yang Liu and Mirella Lapata. 2019 · 2019
Cited alongside, same era.
Towards best experiment design for evaluating dialogue system output
Sashank Santhanam and Samira Shaikh. 2019 · 2019
Cited alongside, same era.
Answers unite! unsupervised metrics for reinforced summarization models
Thomas Scialom, Sylvain Lamprier, Benjamin Piwowarski, and Jacopo Staiano. 2019 · 2019
Cited alongside, same era.
MASS: masked sequence to sequence pre-training for language generation
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2019 · 2019
Cited alongside, same era.
Asking and answering questions to evaluate the factual consistency of summaries
Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020 · 2020
Later among the works it cites.
PEGASUS: pre-training with extracted gap-sentences for abstractive summarization
Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
Summeval: Re-evaluating summarization evaluation
Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Closest in time.
GO FIGURE: A meta evaluation of factuality in summarization
Saadia Gabriel, Asli Celikyilmaz, Rahul Jha, Yejin Choi, and Jianfeng Gao. 2021 · 2021
Closest in time.
The factual inconsistency problem in abstractive text summarization: A survey
Yichong Huang, Xiachong Feng, Xiaocheng Feng, and Bing Qin. 2021 · 2021
Closest in time.
Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics
Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021 · 2021
Closest in time.
How to evaluate a summarizer: Study design and statistical analysis for manual linguistic quality evaluation
Julius Steen and Katja Markert. 2021 · 2021
Closest in time.