Fetching the paper…
Reading the bibliography…
Summarization evaluation remains an open research problem: current metrics such as ROUGE are known to be limited and to correlate poorly with human judgments.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2019 · 1904
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019b · 1904
Earlier work this paper cites.
Question answering as an automatic evaluation metric for news article summarization
Matan Eyal, Tal Baumel, and Michael Elhadad. 2019 · 1906
Earlier work this paper cites.
Answers unite! unsupervised metrics for reinforced summarization models
Thomas Scialom, Sylvain Lamprier, Benjamin Piwowarski, and Jacopo Staiano. 2019 · 1909
Earlier work this paper cites.
Evaluating the factual consistency of abstractive text summarization
Wojciech Kryściński, Bryan McCann, Caiming Xiong, and Richard Socher. 2019 · 1910
Earlier work this paper cites.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019 · 1910
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019 · 1910
Earlier work this paper cites.
Pegasus: Pre-training with extracted gap-sentences for abstractive summarization
Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter J Liu. 2019a · 1912
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Asking and answering questions to evaluate the factual consistency of summaries
Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020 · 2004
Earlier work this paper cites.
Esin Durmus, He He, and Mona Diab. 2020 · 2005
Earlier work this paper cites.
Summeval: Re-evaluating summarization evaluation
Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2020 · 2007
Cited alongside, same era.
METEOR: An automatic metric for MT evaluation with high levels of correlation with human judgments
Alon Lavie and Abhaya Agarwal. 2007 · 2007
Cited alongside, same era.
Reducing quantity hallucinations in abstractive summarization
Zheng Zhao, Shay B Cohen, and Bonnie Webber. 2020 · 2009
Cited alongside, same era.
Discourse constraints for document compression
James Clarke and Mirella Lapata. 2010 · 2010
Cited alongside, same era.
Automatically assessing machine summary content without a gold standard
Annie Louis and Ani Nenkova. 2013 · 2013
Cited alongside, same era.
Neural question generation from text: A preliminary study
Qingyu Zhou, Nan Yang, Furu Wei, Chuanqi Tan, Hangbo Bao, and Ming Zhou. 2017 · 2017
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Improving neural abstractive document summarization with explicit information selection modeling
Wei Li, Xinyan Xiao, Yajuan Lyu, and Yuanzhuo Wang. 2018 · 2018
Later among the works it cites.
Shashi Narayan, Shay B Cohen, and Mirella Lapata. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Cited alongside, same era.
Abstractive text summarization using sequence-to-sequence rnns and beyond
Ramesh Nallapati, Bowen Zhou, Caglar Gulcehre, Bing Xiang, et al. 2016 · 2016
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
Newsqa: A machine comprehension dataset
Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, and Kaheer Suleman. 2016 · 2016
Cited alongside, same era.
A semantic qa-based approach for text summarization evaluation
Ping Chen, Fei Wu, Tong Wang, and Wei Ding. 2017 · 2017
Cited alongside, same era.
Why we need new evaluation metrics for nlg
Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser. 2017 · 2017
Cited alongside, same era.
A deep reinforced model for abstractive summarization
Romain Paulus, Caiming Xiong, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Later among the works it cites.
Object hallucination in image captioning
Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. 2018 · 2018
Later among the works it cites.
Ranking generated summaries by correctness: An interesting but challenging application for natural language inference
Tobias Falke, Leonardo FR Ribeiro, Prasetya Ajie Utama, Ido Dagan, and Iryna Gurevych. 2019 · 2019
Later among the works it cites.
Neural text summarization: A critical evaluation
Wojciech Kryscinski, Nitish Shirish Keskar, Bryan McCann, Caiming Xiong, and Richard Socher. 2019 · 2019
Later among the works it cites.
Studying summarization evaluation metrics in the appropriate scoring range
Maxime Peyrard. 2019 · 2019
Later among the works it cites.
Metrics also disagree in the low scoring range: Revisiting summarization evaluation metrics
Manik Bhandari, Pranav Gour, Atabak Ashfaq, and Pengfei Liu. 2020 · 2020
Later among the works it cites.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020 · 2020
Later among the works it cites.