Fetching the paper…
Reading the bibliography…
Factual consistency is one of important summary evaluation dimensions, especially as summary generation becomes more fluent and coherent.
BERTScore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 1904
Earlier work this paper cites.
How does BERT answer questions? A layer-wise analysis of transformer representations
Betty van Aken, Benjamin Winter, Alexander Löser, and Felix A. Gers. 2019 · 1909
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Automatically evaluating content selection in summarization without human models
Annie Louis and Ani Nenkova. 2009 · 2009
Earlier work this paper cites.
Sensitivity of BLANC to human-scored qualities of text summaries
Oleg Vasilyev, Vedant Dharnidharka, Nicholas Egan, Charlene Chambliss, and John Bohannon. 2020b · 2010
Earlier work this paper cites.
Robust neural abstractive summarization systems and evaluation against adversarial information
Lisa Fan, Dong Yu, and Lu Wang. 2018 · 2018
Earlier work this paper cites.
Ranking generated summaries by correctness: An interesting but challenging application for natural language inference
Tobias Falke, Leonardo FR Ribeiro, Prasetya Ajie Utama, Ido Dagan, and Iryna Gurevych. 2019 · 2019
Earlier work this paper cites.
What does BERT learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019 · 2019
Earlier work this paper cites.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019 · 2019
Earlier work this paper cites.
Answers unite! Unsupervised metrics for reinforced summarization models
Thomas Scialom, Sylvain Lamprier, Benjamin Piwowarski, and Jacopo Staiano. 2019 · 2019
Cited alongside, same era.
The bottom-up evolution of representations in the transformer: A study with machine translation and language modeling objectives
Elena Voita, Rico Sennrich, and Ivan Titov. 2019 · 2019
Cited alongside, same era.
SUM-QE: a BERT-based summary quality estimation model
Stratos Xenouleas, Prodromos Malakasiotis, Marianna Apidianaki, and Ion Androutsopoulos. 2019 · 2019
Cited alongside, same era.
How do decisions emerge across layers in neural models? interpretation with differentiable masking
Nicola De Cao, Michael Sejr Schlichtkrull, Wilker Aziz, and Ivan Titov. 2020 · 2020
Cited alongside, same era.
SUPERT: Towards new frontiers in unsupervised evaluation metrics for multi-document summarization
Yang Gao, Wei Zhao, and Steffen Eger. 2020 · 2020
Cited alongside, same era.
Translation error detection as rationale extraction
Marina Fomicheva, Lucia Specia, and Nikolaos Aletras. 2021 · 2021
Closest in time.
GO FIGURE: A meta evaluation of factuality in summarization
Saadia Gabriel, Asli Celikyilmaz, Rahul Jha, Yejin Choi, and Jianfeng Gao. 2021 · 2021
Closest in time.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021 · 2021
Closest in time.
Analyzing the domain robustness of pretrained language models, layer by layer
Abhinav Ramesh Kashyap, Laiba Mehnaz, Bhavitvya Malik, Abdul Waheed, Devamanyu Hazarika, Min-Yen Kan, and Rajiv Ratn Shah. 2021 · 2021
Closest in time.
QuestEval: Summarization asks for fact-based evaluation
Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari. 2021 · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Evaluating the factual consistency of abstractive text summarization
Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020 · 2020
Cited alongside, same era.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020 · 2020
Cited alongside, same era.
BERTnesia: Investigating the capture and forgetting of knowledge in BERT
Jonas Wallat, Jaspreet Singh, and Avishek Anand. 2020 · 2020
Cited alongside, same era.
Asking and answering questions to evaluate the factual consistency of summaries
Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020 · 2020
Cited alongside, same era.
SummEval: Re-evaluating summarization evaluation
Alexander R. Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021a
Cited in the paper.
QAFactEval: Improved QA-based factual consistency evaluation for summarization
Alexander R. Fabbri, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2021b
Cited in the paper.
ESTIME: Estimation of summary-to-text inconsistency by mismatched embeddings
Oleg Vasilyev and John Bohannon. 2021a
Cited in the paper.
All bark and no bite: Rogue dimensions in transformer language models obscure representational quality
William Timkey and Marten van Schijndel. 2021 · 2021
Closest in time.
Deriving contextualised semantic features from BERT (and other transformer model) embeddings
Jacob Turton, Robert Elliott Smith, and David Vinson. 2021 · 2021
Closest in time.
Is human scoring the best criteria for summary evaluation?
Oleg Vasilyev and John Bohannon. 2021b · 2021
Closest in time.
Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors
Zeyu Yun, Yubei Chen, Bruno Olshausen, and Yann LeCun. 2021 · 2021
Closest in time.