Fetching the paper…
Reading the bibliography…
While neural language models can generate text with remarkable fluency and coherence, controlling for factual correctness in generation remains an open research question.
Unifying human and statistical evaluation for natural language generation
T. Hashimoto, Hugh Zhang, and Percy Liang. 2019 · 1904
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, V. Kishore, Felix Wu, K. Weinberger, and Yoav Artzi. 2020 · 1904
Earlier work this paper cites.
Evaluating the factual consistency of abstractive text summarization
Wojciech Kryscinski, B. McCann, Caiming Xiong, and R. Socher. 2019 · 1910
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, W. Li, and Peter J. Liu. 2019 · 1910
Earlier work this paper cites.
Samsum corpus: A human-annotated dialogue dataset for abstractive summarization
Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Aleksander Wawer. 2019 · 1911
Earlier work this paper cites.
How decoding strategies affect the verifiability of generated text
Luca Massarelli, F. Petroni, Aleksandra Piktus, Myle Ott, Tim Rocktäschel, Vassilis Plachouras, F. Silvestri, and S. Riedel. 2019 · 1911
Earlier work this paper cites.
The measurement of observer agreement for categorical data
J. Landis and G. Koch. 1977 · 1977
Earlier work this paper cites.
Wordnet: A lexical database for english
George A. Miller. 1995 · 1995
Earlier work this paper cites.
Fill in the blanc: Human-free quality estimation of document summaries
Oleg V. Vasilyev, Vedant Dharnidharka, and J. Bohannon. 2020 · 2002
Earlier work this paper cites.
Boosting factual correctness of abstractive summarization
Chenguang Zhu, William Hinthorn, Ruochen Xu, Qingkai Zeng, Michael Zeng, Xuedong Huang, and Meng Jiang. 2020 · 2003
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Earlier work this paper cites.
(meta-) evaluation of machine translation
Chris Callison-Burch, Cameron Fordyce, Philipp Koehn, Christof Monz, and Josh Schroeder. 2007 · 2007
Earlier work this paper cites.
An investigation into the validity of some metrics for automatically evaluating natural language generation systems
Ehud Reiter and Anja Belz. 2009 · 2009
Cited alongside, same era.
Reducing quantity hallucinations in abstractive summarization
Z. Zhao, Shay B. Cohen, and B. Webber. 2020 · 2009
Cited alongside, same era.
A decade of automatic content evaluation of news summaries: Reassessing the state of the art
Peter A. Rankel, John M. Conroy, Hoa Trang Dang, and Ani Nenkova. 2013 · 2013
Cited alongside, same era.
Evaluating machine translation for assimilation via a gap-filling task
E. Ageeva, M. Forcada, Francis M. Tyers, and Juan Antonio Pérez-Ortiz. 2015 · 2015
Cited alongside, same era.
Re-evaluating automatic summarization with BLEU and 192 shades of ROUGE
Yvette Graham. 2015 · 2015
Cited alongside, same era.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern. 2018 · 2018
Later among the works it cites.
Sentence mover’s similarity: Automatic evaluation for multi-sentence texts
Elizabeth Clark, Asli Celikyilmaz, and Noah A. Smith. 2019 · 2019
Later among the works it cites.
Question answering as an automatic evaluation metric for news article summarization
Matan Eyal, Tal Baumel, and Michael Elhadad. 2019 · 2019
Later among the works it cites.
Assessing the factual accuracy of generated text
B. Goodrich, V. Rao, Mohammad Saleh, and Peter J. Liu. 2019 · 2019
Later among the works it cites.
Hierarchical transformers for multi-document summarization
Yang Liu and Mirella Lapata. 2019 · 2019
Later among the works it cites.
Answers unite! unsupervised metrics for reinforced summarization models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Abstractive text summarization using sequence-to-sequence rnns and beyond
Ramesh Nallapati, Bowen Zhou, C. D. Santos, Çaglar Gülçehre, and B. Xiang. 2016 · 2016
Cited alongside, same era.
Squad: 100, 000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
Why we need new evaluation metrics for nlg
Jekaterina Novikova, Ondrej Dusek, A. Curry, and Verena Rieser. 2017 · 2017
Cited alongside, same era.
Faithful to the original: Fact aware neural abstractive summarization
Ziqiang Cao, Furu Wei, W. Li, and Sujian Li. 2018 · 2018
Cited alongside, same era.
The price of debiasing automatic metrics in natural language evalaution
Arun Chaganty, Stephen Mussmann, and Percy Liang. 2018 · 2018
Cited alongside, same era.
Soft layer-specific multi-task summarization with entailment and question generation
Han Guo, Ramakanth Pasunuru, and Mohit Bansal. 2018 · 2018
Cited alongside, same era.
Don’t give me the details, just the summary! Topic-aware convolutional neural networks for extreme summarization
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018 · 2018
Cited alongside, same era.
Thomas Scialom, Sylvain Lamprier, Benjamin Piwowarski, and Jacopo Staiano. 2019 · 2019
Later among the works it cites.
FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization
Esin Durmus, He He, and Mona Diab. 2020 · 2020
Closest in time.
Knowledge graph-augmented abstractive summarization with semantic-driven cloze reward
Luyang Huang, Lingfei Wu, and Lu Wang. 2020 · 2020
Closest in time.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Closest in time.
Unsupervised commonsense question answering with self-talk
Vered Shwartz, Peter West, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020 · 2020
Closest in time.
Asking and answering questions to evaluate the factual consistency of summaries
Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020 · 2020
Closest in time.
Summeval: Re-evaluating summarization evaluation
A. R. Fabbri, Wojciech Kryscinski, Bryan McCann, R. Socher, and Dragomir Radev. 2021 · 2021
Closest in time.