Fetching the paper…
Reading the bibliography…
Recent advancements in text summarization, particularly with the advent of Large Language Models (LLMs), have shown remarkable performance.
Adversarial NLI: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2019 · 1910
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019 · 1910
Earlier work this paper cites.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020 · 1919
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Automation of summary evaluation by the pyramid method
Aaron Harnly, Ani Nenkova, Rebecca Passonneau, and Owen Rambow. 2005 · 2005
Earlier work this paper cites.
Deberta: Decoding-enhanced BERT with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2020 · 2006
Earlier work this paper cites.
The pyramid method: Incorporating human content selection variation in summarization evaluation
Ani Nenkova, Rebecca Passonneau, and Kathleen McKeown. 2007 · 2007
Earlier work this paper cites.
The balanced accuracy and its posterior distribution
Kay Henning Brodersen, Cheng Soon Ong, Klaas Enno Stephan, and Joachim M. Buhmann. 2010 · 2010
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence rnns and beyond
Ramesh Nallapati, Bowen Zhou, Cicero Nogueira dos santos, Caglar Gulcehre, and Bing Xiang. 2016 · 2016
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2017 · 2017
Earlier work this paper cites.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018 · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern. 2018 · 2018
Earlier work this paper cites.
Automated pyramid summarization evaluation
Yanjun Gao, Chen Sun, and Rebecca J. Passonneau. 2019 · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Earlier work this paper cites.
FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization
Esin Durmus, He He, and Mona Diab. 2020 · 2020
Earlier work this paper cites.
Evaluating factuality in generation with dependency-level entailment
Tanya Goyal and Greg Durrett. 2020 · 2020
Cited alongside, same era.
Evaluating the factual consistency of abstractive text summarization
Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020 · 2020
Cited alongside, same era.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Cited alongside, same era.
Asking and answering questions to evaluate the factual consistency of summaries
Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020 · 2020
Cited alongside, same era.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Cited alongside, same era.
F-coref: Fast, accurate and easy to use coreference resolution
Shon Otmazgin, Arie Cattan, and Yoav Goldberg. 2022 · 2022
Later among the works it cites.
FactGraph: Evaluating factuality in summarization with semantic graph representations
Leonardo F. R. Ribeiro, Mengwen Liu, Iryna Gurevych, Markus Dreyer, and Mohit Bansal. 2022 · 2022
Later among the works it cites.
Evaluating the factual consistency of large language models through summarization
Derek Tam, Anisha Mascarenhas, Shiyue Zhang, Sarah Kwan, Mohit Bansal, and Colin Raffel. 2022 · 2022
Later among the works it cites.
Evaluating factual consistency of summaries with large language models
Shiqi Chen, Siyang Gao, and Junxian He. 2023 · 2023
Later among the works it cites.
Menli: Robust evaluation metrics from natural language inference
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
SummEval: Re-evaluating summarization evaluation
Alexander R. Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Cited alongside, same era.
Annotating and modeling fine-grained factuality in summarization
Tanya Goyal and Greg Durrett. 2021 · 2021
Cited alongside, same era.
SumPubMed: Summarization dataset of PubMed scientific articles
Vivek Gupta, Prerna Bharti, Pegah Nokhiz, and Harish Karnick. 2021 · 2021
Cited alongside, same era.
Does putting a linguist in the loop improve NLU data collection?
Alicia Parrish, William Huang, Omar Agha, Soo-Hwan Lee, Nikita Nangia, Alexia Warstadt, Karmanya Aggarwal, Emily Allaway, Tal Linzen, and Samuel R. Bowman. 2021 · 2021
Cited alongside, same era.
QuestEval: Summarization asks for fact-based evaluation
Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari. 2021 · 2021
Cited alongside, same era.
Finding a balanced degree of automation for summary evaluation
Shiyue Zhang and Mohit Bansal. 2021 · 2021
Cited alongside, same era.
QAFactEval: Improved QA-based factual consistency evaluation for summarization
Alexander Fabbri, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2022 · 2022
Cited alongside, same era.
Yanran Chen and Steffen Eger. 2023 · 2023
Later among the works it cites.
Evaluating factual consistency of texts with semantic role labeling
Jing Fan, Dennis Aumiller, and Michael Gertz. 2023 · 2023
Later among the works it cites.
Gptscore: Evaluate as you desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023 · 2023
Later among the works it cites.
Human-like summarization evaluation with chatgpt
Mingqi Gao, Jie Ruan, Renliang Sun, Xunjian Yin, Shiping Yang, and Xiaojun Wan. 2023 · 2023
Later among the works it cites.
Trueteacher: Learning factual consistency evaluation with large language models
Zorik Gekhman, Jonathan Herzig, Roee Aharoni, Chen Elkind, and Idan Szpektor. 2023 · 2023
Later among the works it cites.
LongEval: Guidelines for human evaluation of faithfulness in long-form summarization
Kalpesh Krishna, Erin Bransom, Bailey Kuehl, Mohit Iyyer, Pradeep Dasigi, Arman Cohan, and Kyle Lo. 2023 · 2023
Later among the works it cites.
Chatgpt as a factual inconsistency evaluator for text summarization
Zheheng Luo, Qianqian Xie, and Sophia Ananiadou. 2023 · 2023
Later among the works it cites.
Factscore: Fine-grained atomic evaluation of factual precision in long form text generation
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023 · 2023
Later among the works it cites.
Echoes from alexandria: A large resource for multilingual book summarization
Alessandro Scirè, Simone Conia, Simone Ciciliano, and Roberto Navigli. 2023 · 2023
Later among the works it cites.
Large language models are not yet human-level evaluators for abstractive summarization
Chenhui Shen, Liying Cheng, Xuan-Phi Nguyen, Yang You, and Lidong Bing. 2023 · 2023
Later among the works it cites.
Understanding factual errors in summarization: Errors, summarizers, datasets, error detectors
Liyan Tang, Tanya Goyal, Alex Fabbri, Philippe Laban, Jiacheng Xu, Semih Yavuz, Wojciech Kryscinski, Justin Rousseau, and Greg Durrett. 2023 · 2023
Later among the works it cites.
AlignScore: Evaluating factual consistency with a unified alignment function
Yuheng Zha, Yichi Yang, Ruichen Li, and Zhiting Hu. 2023 · 2023
Later among the works it cites.