Fetching the paper…
Reading the bibliography…
Factual consistency evaluation is often conducted using Natural Language Inference (NLI) models, yet these models exhibit limited success in evaluating summaries.
Assessing the factual accuracy of generated text
Ben Goodrich, Vinay Rao, Mohammad Saleh, and Peter J. Liu. 2019 · 1905
Earlier work this paper cites.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan T. McDonald. 2020 · 1919
Earlier work this paper cites.
Open information extraction from the web
Michele Banko, Michael J. Cafarella, Stephen Soderland, Matthew Broadhead, and Oren Etzioni. 2007 · 2007
Earlier work this paper cites.
Summeval: Re-evaluating summarization evaluation
Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2020 · 2007
Earlier work this paper cites.
Wikilingua: A new benchmark dataset for cross-lingual abstractive summarization
Faisal Ladhak, Esin Durmus, Claire Cardie, and Kathleen R. McKeown. 2020 · 2010
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomás Kociský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
XNLI: evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel R. Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Earlier work this paper cites.
Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies
Max Grusky, Mor Naaman, and Yoav Artzi. 2018 · 2018
Earlier work this paper cites.
Scitail: A textual entailment dataset from science question answering
Tushar Khot, Ashish Sabharwal, and Peter Clark. 2018 · 2018
Earlier work this paper cites.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018 · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Ranking generated summaries by correctness: An interesting but challenging application for natural language inference
Tobias Falke, Leonardo F. R. Ribeiro, Prasetya Ajie Utama, Ido Dagan, and Iryna Gurevych. 2019a · 2019
Earlier work this paper cites.
Neural text summarization: A critical evaluation
Wojciech Kryscinski, Nitish Shirish Keskar, Bryan McCann, Caiming Xiong, and Richard Socher. 2019 · 2019
Earlier work this paper cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020 · 2020
Earlier work this paper cites.
Evaluating the factual consistency of abstractive text summarization
Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020 · 2020
Earlier work this paper cites.
BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Earlier work this paper cites.
Adversarial NLI: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Earlier work this paper cites.
Asking and answering questions to evaluate the factual consistency of summaries
Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020 · 2020
Earlier work this paper cites.
Xl-sum: Large-scale multilingual abstractive summarization for 44 languages
Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Samin Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021 · 2021
Earlier work this paper cites.
$qˆ2$: Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering
Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend. 2021 · 2021
Cited alongside, same era.
Looking beyond sentence-level natural language inference for question answering and text summarization
Anshuman Mishra, Dhruvesh Patel, Aparna Vijayakumar, Xiang Lorraine Li, Pavan Kapanipathi, and Kartik Talamadupula. 2021 · 2021
Cited alongside, same era.
Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics
Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021 · 2021
Cited alongside, same era.
Measuring attribution in natural language generation models
Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Michael Collins, Dipanjan Das, Slav Petrov, Gaurav Singh Tomar, Iulia Turc, and David Reitter. 2021 · 2021
Cited alongside, same era.
Questeval: Summarization asks for fact-based evaluation
WANLI: worker and AI collaboration for natural language inference dataset creation
Alisa Liu, Swabha Swayamdipta, Noah A. Smith, and Yejin Choi. 2022 · 2022
Later among the works it cites.
Chatgpt, https://openai.com/blog/chatgpt/
OpenAI. 2022 · 2022
Later among the works it cites.
Stretching sentence-pair NLI models to reason over long documents and clusters
Tal Schuster, Sihao Chen, Senaka Buthpitiya, Alex Fabrikant, and Donald Metzler. 2022 · 2022
Later among the works it cites.
Falsesum: Generating document-level NLI examples for recognizing factual inconsistency in summarization
Prasetya Utama, Joshua Bambrick, Nafise Sadat Moosavi, and Iryna Gurevych. 2022 · 2022
Later among the works it cites.
Symbolic knowledge distillation: from general language models to commonsense models
Peter West, Chandra Bhagavatula, Jack Hessel, Jena Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, and Yejin Choi. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari. 2021 · 2021
Cited alongside, same era.
Want to reduce labeling cost? GPT-3 can help
Shuohang Wang, Yang Liu, Yichong Xu, Chenguang Zhu, and Michael Zeng. 2021 · 2021
Cited alongside, same era.
mt5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2021
Cited alongside, same era.
Docnli: A large-scale dataset for document-level natural language inference
Wenpeng Yin, Dragomir R. Radev, and Caiming Xiong. 2021 · 2021
Cited alongside, same era.
Qameleon: Multilingual QA with only 5 examples
Priyanka Agrawal, Chris Alberti, Fantine Huot, Joshua Maynez, Ji Ma, Sebastian Ruder, Kuzman Ganchev, Dipanjan Das, and Mirella Lapata. 2022 · 2022
Cited alongside, same era.
mface: Multilingual summarization with factual consistency evaluation
Roee Aharoni, Shashi Narayan, Joshua Maynez, Jonathan Herzig, Elizabeth Clark, and Mirella Lapata. 2022 · 2022
Cited alongside, same era.
Correcting diverse factual errors in abstractive summarization via post-editing and language model infilling
Vidhisha Balachandran, Hannaneh Hajishirzi, William W. Cohen, and Yulia Tsvetkov. 2022 · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022 · 2022
Cited alongside, same era.
Rachith Aiyappa, Jisun An, Haewoon Kwak, and Yong-Yeol Ahn. 2023 · 2023
Closest in time.
q2d: Turning questions into dialogs to teach models how to search
Yonatan Bitton, Shlomi Cohen-Ganor, Ido Hakimi, Yoad Lewenberg, Roee Aharoni, and Enav Weinreb. 2023 · 2023
Closest in time.
Evaluating factual consistency of summaries with large language models
Shiqi Chen, Siyang Gao, and Junxian He. 2023 · 2023
Closest in time.
Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes
Cheng-Yu Hsieh, Chun-Liang Li, Chih-kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alex Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister. 2023 · 2023
Closest in time.
Large language models are state-of-the-art evaluators of translation quality
Tom Kocmi and Christian Federmann. 2023 · 2023
Closest in time.
An empirical survey on long document summarization: Datasets, models, and metrics
Huan Yee Koh, Jiaxin Ju, Ming Liu, and Shirui Pan. 2023 · 2023
Closest in time.
G-eval: NLG evaluation using GPT-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Closest in time.
Chatgpt as a factual inconsistency evaluator for abstractive text summarization
Zheheng Luo, Qianqian Xie, and Sophia Ananiadou. 2023 · 2023
Closest in time.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Sean Welleck, Bodhisattwa Prasad Majumder, Shashank Gupta, Amir Yazdanbakhsh, and Peter Clark. 2023 · 2023
Closest in time.
NonFactS: NonFactual summary generation for factuality evaluation in document summarization
Amir Soleimani, Christof Monz, and Marcel Worring. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023 · 2023
Closest in time.
Is chatgpt a good NLG evaluator? A preliminary study
Jiaan Wang, Yunlong Liang, Fandong Meng, Haoxiang Shi, Zhixu Li, Jinan Xu, Jianfeng Qu, and Jie Zhou. 2023 · 2023
Closest in time.
Large language models are better reasoners with self-verification
Yixuan Weng, Minjun Zhu, Fei Xia, Bin Li, Shizhu He, Kang Liu, and Jun Zhao. 2023 · 2023
Closest in time.
Wecheck: Strong factual consistency checker via weakly supervised learning
Wenhao Wu, Wei Li, Xinyan Xiao, Jiachen Liu, Sujian Li, and Yajuan Lv. 2023 · 2023
Closest in time.