GO FIGURE: A meta evaluation of factuality in summarization
Saadia Gabriel, Asli Celikyilmaz, Rahul Jha, Yejin Choi, and Jianfeng Gao. 2021 · 2021
Later among the works it cites.
The GEM benchmark: Natural language generation, its evaluation and metrics
Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, Wanyu Du, Esin Durmus, Ondřej Dušek, Chris Chinenye Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Mihir Kale, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Andre Niyongabo Rubungo, Salomey Osei, Ankur Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou. 2021 · 2021
Later among the works it cites.
Dialfact: A benchmark for fact-checking in dialogue
Original
Prakhar Gupta, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2021 · 2021
Later among the works it cites.
Deberta: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021 · 2021
Later among the works it cites.
q 2 q^{2} : Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering
Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend. 2021 · 2021
Later among the works it cites.
Summac: Re-visiting nli-based models for inconsistency detection in summarization
Original
Philippe Laban, Tobias Schnabel, Paul N Bennett, and Marti A Hearst. 2021 · 2021
Later among the works it cites.
I like fish, especially dolphins: Addressing contradictions in dialogue modeling
Yixin Nie, Mary Williamson, Mohit Bansal, Douwe Kiela, and Jason Weston. 2021 · 2021
Later among the works it cites.
Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics
Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021 · 2021
Later among the works it cites.
Quality: Question answering with long input texts, yes!
Original
Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He, et al. 2021 · 2021
Later among the works it cites.
Learning compact metrics for MT
Amy Pu, Hyung Won Chung, Ankur Parikh, Sebastian Gehrmann, and Thibault Sellam. 2021 · 2021
Later among the works it cites.
Don’t be contradicted with anything! CI-ToD: Towards benchmarking consistency for task-oriented dialogue system
Libo Qin, Tianbao Xie, Shijue Huang, Qiguang Chen, Xiao Xu, and Wanxiang Che. 2021 · 2021
Later among the works it cites.
Get your vitamin C! robust fact verification with contrastive evidence
Tal Schuster, Adam Fisch, and Regina Barzilay. 2021 · 2021
Later among the works it cites.
QuestEval: Summarization asks for fact-based evaluation
Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari. 2021 · 2021
Later among the works it cites.
Beametrics: A benchmark for language generation evaluation evaluation
Original
Thomas Scialom and Felix Hill. 2021 · 2021
Later among the works it cites.
Factual consistency evaluation for text summarization via counterfactual estimation
Yuexiang Xie, Fei Sun, Yang Deng, Yaliang Li, and Bolin Ding. 2021 · 2021
Later among the works it cites.
A comprehensive assessment of dialog evaluation metrics
Yi-Ting Yeh, Maxine Eskenazi, and Shikib Mehri. 2021 · 2021
Later among the works it cites.
DocNLI: A large-scale dataset for document-level natural language inference
Wenpeng Yin, Dragomir Radev, and Caiming Xiong. 2021 · 2021
Later among the works it cites.
BARTScore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Later among the works it cites.
Detecting hallucinated content in conditional neural sequence generation
Chunting Zhou, Graham Neubig, Jiatao Gu, Mona Diab, Francisco Guzmán, Luke Zettlemoyer, and Marjan Ghazvininejad. 2021 · 2021
Later among the works it cites.
Re-examining system-level correlations of automatic summarization evaluation metrics
Original
Daniel Deutsch, Rotem Dror, and Dan Roth. 2022 · 2022
Closest in time.
Scrolls: Standardized comparison over long language sequences
Original
Uri Shaham, Elad Segal, Maor Ivgi, Avia Efrat, Ori Yoran, Adi Haviv, Ankit Gupta, Wenhan Xiong, Mor Geva, Jonathan Berant, and Omer Levy. 2022 · 2022
Closest in time.
Are factuality checkers reliable? adversarial meta-evaluation of factuality in summarization
Yiran Chen, Pengfei Liu, and Xipeng Qiu. 2021 · 2095
Closest in time.