Fetching the paper…
Reading the bibliography…
Automated evaluation of text generation systems has recently seen increasing attention, particularly checking whether generated text stays truthful to input sources.
Simple BERT models for relation extraction and semantic role labeling
Peng Shi and Jimmy Lin. 2019 · 1904
Earlier work this paper cites.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020 · 1919
Earlier work this paper cites.
Multiple comparisons among means
Olive Jean Dunn. 1961 · 1961
Earlier work this paper cites.
Functionalist approaches to grammar
Elizabeth Bates and Brian Macwhinney. 1982 · 1982
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
The Proposition Bank: An annotated corpus of semantic roles
Martha Palmer, Daniel Gildea, and Paul Kingsbury. 2005 · 2005
Earlier work this paper cites.
On some pitfalls in automatic evaluation and significance testing for MT
Stefan Riezler and John T. Maxwell. 2005 · 2005
Earlier work this paper cites.
Can semantic roles generalize across genres?
Szu-ting Yi, Edward Loper, and Martha Palmer. 2007 · 2007
Earlier work this paper cites.
Semantic role labeling: An introduction to the special issue
Lluís Màrquez, Xavier Carreras, Kenneth C. Litkowski, and Suzanne Stevenson. 2008 · 2008
Earlier work this paper cites.
Speech and language processing: an introduction to natural language processing, computational linguistics, and speech recognition, 2nd Edition
Dan Jurafsky and James H. Martin. 2009 · 2009
Earlier work this paper cites.
VerbNet class assignment as a WSD task
Susan Windisch Brown, Dmitriy Dligach, and Martha Palmer. 2011 · 2011
Earlier work this paper cites.
English propbank annotation guidelines
Claire Bonial, Jena Hwang, Julia Bonn, Kathryn Conger, Olga Babko-Malaya, and Martha Palmer. 2012 · 2012
Earlier work this paper cites.
What makes a good summary? reconsidering the focus of automatic summarization
Maartje ter Hoeve, Julia Kiseleva, and Maarten de Rijke. 2020 · 2012
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomás Kociský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Cited alongside, same era.
Abstractive text summarization using sequence-to-sequence rnns and beyond
Ramesh Nallapati, Bowen Zhou, Cícero Nogueira dos Santos, Çaglar Gülçehre, and Bing Xiang. 2016 · 2016
Cited alongside, same era.
End-to-end neural coreference resolution
Kenton Lee, Luheng He, Mike Lewis, and Luke Zettlemoyer. 2017 · 2017
Cited alongside, same era.
Summarunner: A recurrent neural network based sequence model for extractive summarization of documents
Ramesh Nallapati, Feifei Zhai, and Bowen Zhou. 2017 · 2017
Cited alongside, same era.
Get to the point: Summarization with pointer-generator networks
Abigail See, Peter J. Liu, and Christopher D. Manning. 2017 · 2017
Cited alongside, same era.
The hitchhiker’s guide to testing statistical significance in natural language processing
Asking and answering questions to evaluate the factual consistency of summaries
Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020 · 2020
Later among the works it cites.
Bertscore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Later among the works it cites.
A statistical analysis of summarization evaluation metrics using resampling methods
Daniel Deutsch, Rotem Dror, and Dan Roth. 2021 · 2021
Later among the works it cites.
SummEval: Re-evaluating summarization evaluation
Alexander R. Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Later among the works it cites.
Looking beyond sentence-level natural language inference for question answering and text summarization
Anshuman Mishra, Dhruvesh Patel, Aparna Vijayakumar, Xiang Lorraine Li, Pavan Kapanipathi, and Kartik Talamadupula. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rotem Dror, Gili Baumer, Segev Shlomov, and Roi Reichart. 2018 · 2018
Cited alongside, same era.
AllenNLP: A deep semantic natural language processing platform
Matt Gardner, Joel Grus, Mark Neumann, Oyvind Tafjord, Pradeep Dasigi, Nelson F. Liu, Matthew Peters, Michael Schmitz, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Bottom-up abstractive summarization
Sebastian Gehrmann, Yuntian Deng, and Alexander Rush. 2018 · 2018
Cited alongside, same era.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018 · 2018
Cited alongside, same era.
Ranking generated summaries by correctness: An interesting but challenging application for natural language inference
Tobias Falke, Leonardo F. R. Ribeiro, Prasetya Ajie Utama, Ido Dagan, and Iryna Gurevych. 2019 · 2019
Cited alongside, same era.
Assessing the factual accuracy of generated text
Ben Goodrich, Vinay Rao, Peter J. Liu, and Mohammad Saleh. 2019 · 2019
Cited alongside, same era.
FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization
Esin Durmus, He He, and Mona Diab. 2020 · 2020
Cited alongside, same era.
Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics
Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021 · 2021
Later among the works it cites.
Factual consistency evaluation for text summarization via counterfactual estimation
Yuexiang Xie, Fei Sun, Yang Deng, Yaliang Li, and Bolin Ding. 2021 · 2021
Later among the works it cites.
Conversational semantic role labeling
Kun Xu, Han Wu, Linfeng Song, Haisong Zhang, Linqi Song, and Dong Yu. 2021 · 2021
Later among the works it cites.
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Later among the works it cites.
Semantic representations for NLP using verbnet and the generative lexicon
Susan Windisch Brown, Julia Bonn, Ghazaleh Kazeminejad, Annie Zaenen, James Pustejovsky, and Martha Palmer. 2022 · 2022
Later among the works it cites.
Re-examining system-level correlations of automatic summarization evaluation metrics
Daniel Deutsch, Rotem Dror, and Dan Roth. 2022 · 2022
Later among the works it cites.
Measuring faithfulness of abstractive summaries
Tim Fischer, Steffen Remus, and Chris Biemann. 2022 · 2022
Later among the works it cites.
Yiyang Li, Lei Li, Qing Yang, Marina Litvak, Natalia Vanetik, Dingxin Hu, Yuze Li, Yanquan Zhou, Dongliang Xu, and Xuanyu Zhang. 2022 · 2022
Later among the works it cites.
PropBank comes of Age—Larger, smarter, and more diverse
Sameer Pradhan, Julia Bonn, Skatje Myers, Kathryn Conger, Tim O’gorman, James Gung, Kristin Wright-bettner, and Martha Palmer. 2022 · 2022
Later among the works it cites.
G-eval: NLG evaluation using GPT-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Closest in time.