Fetching the paper…
Reading the bibliography…
Canonical automatic summary evaluation metrics, such as ROUGE, focus on lexical similarity which cannot well capture semantics nor linguistic quality and require a reference summary which is costly to obtain.
Better summarization evaluation with word embeddings for rouge
Jun-Ping Ng and Viktoria Abrecht. 2015 · 1930
Earlier work this paper cites.
Automatic dialogue summary generation for customer service
Chunyi Liu, Peng Wang, Jiang Xu, Zang Li, and Jieping Ye. 2019 · 1965
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Correlation between rouge and human evaluation of extractive meeting summaries
Feifan Liu and Yang Liu. 2008 · 2008
Earlier work this paper cites.
Automatic summarization of bug reports
S. Rastkar, G. C. Murphy, and G. Murray. 2014 · 2014
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
How not to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation
Chia-Wei Liu, Ryan Lowe, Iulian Serban, Mike Noseworthy, Laurent Charlin, and Joelle Pineau. 2016 · 2016
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence rnns and beyond
Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Caglar Gulcehre, and Bing Xiang. 2016 · 2016
Earlier work this paper cites.
Machine learning approach to evaluate multilingual summaries
Samira Ellouze, Maher Jaoua, and Lamia Hadrich Belguith. 2017 · 2017
Earlier work this paper cites.
Learning to score system summaries for better content selection evaluation
Maxime Peyrard, Teresa Botschen, and Iryna Gurevych. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
A discourse-aware attention model for abstractive summarization of long documents
Arman Cohan, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Seokhwan Kim, Walter Chang, and Nazli Goharian. 2018 · 2018
Cited alongside, same era.
Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies
Max Grusky, Mor Naaman, and Yoav Artzi. 2018 · 2018
Cited alongside, same era.
Scoring summaries using recurrent neural networks
Stefan Ruseti, Mihai Dascalu, Amy M Johnson, Danielle S McNamara, Renu Balyan, Kathryn S McCarthy, and Stefan Trausan-Matu. 2018 · 2018
Cited alongside, same era.
Answers unite! unsupervised metrics for reinforced summarization models
Thomas Scialom, Sylvain Lamprier, Benjamin Piwowarski, and Jacopo Staiano. 2019 · 2019
Later among the works it cites.
BIGPATENT: A large-scale dataset for abstractive and coherent summarization
Eva Sharma, Chen Li, and Lu Wang. 2019 · 2019
Later among the works it cites.
Automatic learner summary assessment for reading comprehension
Menglin Xia, Ekaterina Kochmar, and Ted Briscoe. 2019 · 2019
Later among the works it cites.
MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019 · 2019
Later among the works it cites.
Re-evaluating evaluation in text summarization
Manik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu, and Graham Neubig. 2020 · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Guokan Shang, Wensi Ding, Zekun Zhang, Antoine Tixier, Polykarpos Meladianos, Michalis Vazirgiannis, and Jean-Pierre Lorré. 2018 · 2018
Cited alongside, same era.
Making sense of group chat through collaborative tagging and summarization
Amy X. Zhang and Justin Cranshaw. 2018 · 2018
Cited alongside, same era.
Estimating summary quality with pairwise preferences
Markus Zopf. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
BillSum: A corpus for automatic summarization of US legislation
Anastassia Kornilova and Vladimir Eidelman. 2019 · 2019
Cited alongside, same era.
Yang Gao, Wei Zhao, and Steffen Eger. 2020 · 2020
Closest in time.
Fill in the BLANC: Human-free quality estimation of document summaries
Oleg Vasilyev, Vedant Dharnidharka, and John Bohannon. 2020 · 2020
Closest in time.
Unsupervised reference-free summary quality evaluation via contrastive learning
Hanlu Wu, Tengfei Ma, Lingfei Wu, Tariro Manyumwa, and Shouling Ji. 2020 · 2020
Closest in time.
BERTScore: Evaluating text generation with BERT
Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Closest in time.
TAC2010 guided summarization competition
NIST. 2010 · 2021
Closest in time.