Fetching the paper…
Reading the bibliography…
With recent improvements in natural language generation (NLG) models for various applications, it has become imperative to have the means to identify and evaluate whether NLG output is only sharing verifiable information about the external world.
Logic and conversation
Grice, Herbert P. 1975 · 1975
Earlier work this paper cites.
Harvey Friedman’s research on the foundations of mathematics
Harrington, Leo A, Michael D Morley, A Šcedrov, and Stephen G Simpson. 1985 · 1985
Earlier work this paper cites.
Implicature, explicature, and truth-theoretic semantics
Carston, Robyn. 1988 · 1988
Earlier work this paper cites.
Why the child’s theory of mind really is a theory
Gopnik, Alison and Henry M. Wellman. 1992 · 1992
Earlier work this paper cites.
Towards a human-like open-domain chatbot
Adiwardana, Daniel, Minh-Thang Luong, David R So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, et al. 2020 · 2001
Earlier work this paper cites.
Relevance Theory , The Handbook of Pragmatics. Blackwell
Wilson, Deirdre and Dan Sperber. 2004 · 2004
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence RNNs and beyond
Nallapati, Ramesh, Bowen Zhou, Cicero dos Santos, Çağlar Gu̇lçehre, and Bing Xiang. 2016 · 2016
Earlier work this paper cites.
Model-agnostic interpretability of machine learning
Ribeiro, Marco Tulio, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Get to the point: Summarization with pointer-generator networks
See, Abigail, Peter J. Liu, and Christopher D. Manning. 2017 · 2017
Earlier work this paper cites.
Challenges in data-to-document generation
Wiseman, Sam, Stuart Shieber, and Alexander Rush. 2017 · 2017
Earlier work this paper cites.
QuAC: Question answering in context
Choi, Eunsol, He He, Mohit Iyyer, Mark Yatskar, Wen-tau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer. 2018 · 2018
Earlier work this paper cites.
Automated fact checking: Task formulations, methods and future directions
Thorne, James and Andreas Vlachos. 2018 · 2018
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
Thorne, James, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018 · 2018
Earlier work this paper cites.
Wizard of wikipedia: Knowledge-powered conversational agents
Dinan, Emily, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2019 · 2019
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
Kwiatkowski, Tom, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Cited alongside, same era.
Combining fact extraction and verification with neural semantic matching networks
Nie, Yixin, Haonan Chen, and Mohit Bansal. 2019 · 2019
Cited alongside, same era.
Inherent Disagreements in Human Textual Inferences
Pavlick, Ellie and Tom Kwiatkowski. 2019 · 2019
Cited alongside, same era.
Dialogue natural language inference
Welleck, Sean, Jason Weston, Arthur D. Szlam, and Kyunghyun Cho. 2019 · 2019
Cited alongside, same era.
Disentangling the properties of human evaluation methods: A classification system to support comparability, meta-evaluation and reproducibility testing
Belz, Anja, Simon Mille, and David M Howcroft. 2020 · 2020
Cited alongside, same era.
Big bird: Transformers for longer sequences
Zaheer, Manzil, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. 2020 · 2020
Later among the works it cites.
Extractive summarization as text matching
Zhong, Ming, Pengfei Liu, Yiran Chen, Danqing Wang, Xipeng Qiu, and Xuanjing Huang. 2020 · 2020
Later among the works it cites.
Open-domain question answering goes conversational via question rewriting
Anantha, Raviteja, Svitlana Vakulenko, Zhucheng Tu, Shayne Longpre, Stephen Pulman, and Srinivas Chappidi. 2021 · 2021
Closest in time.
Decontextualization: Making Sentences Stand-Alone
Choi, Eunsol, Jennimaria Palomaki, Matthew Lamm, Tom Kwiatkowski, Dipanjan Das, and Michael Collins. 2021 · 2021
Closest in time.
Evaluating groundedness in dialogue systems: The BEGIN benchmark
Dziri, Nouha, Hannah Rashkin, Tal Linzen, and David Reitter. 2021 · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cast-19: A dataset for conversational information seeking
Dalton, Jeffrey, Chenyan Xiong, Vaibhav Kumar, and Jamie Callan. 2020 · 2020
Cited alongside, same era.
FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization
Durmus, Esin, He He, and Mona Diab. 2020 · 2020
Cited alongside, same era.
Twenty years of confusion in human evaluation: Nlg needs evaluation sheets and standardised definitions
Howcroft, David M, Anja Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, and Verena Rieser. 2020 · 2020
Cited alongside, same era.
On faithfulness and factuality in abstractive summarization
Maynez, Joshua, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020 · 2020
Cited alongside, same era.
USR: An unsupervised and reference free evaluation metric for dialog generation
Mehri, Shikib and Maxine Eskenazi. 2020 · 2020
Cited alongside, same era.
ToTTo: A controlled table-to-text generation dataset
Parikh, Ankur, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, and Dipanjan Das. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, Colin, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Cited alongside, same era.
The GEM benchmark: Natural language generation, its evaluation and metrics
Gehrmann, Sebastian, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, Wanyu Du, Esin Durmus, Ondřej Dušek, Chris Chinenye Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Mihir Kale, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Andre Niyongabo Rubungo, Salomey Osei, Ankur Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou. 2021 · 2021
Closest in time.
Dialfact: A benchmark for fact-checking in dialogue
Gupta, Prakhar, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2021 · 2021
Closest in time.
q 2 q^{2} : Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering
Honovich, Or, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend. 2021 · 2021
Closest in time.
Improving factual consistency of abstractive summarization via question answering
Nan, Feng, Cicero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Kathleen McKeown, Ramesh Nallapati, Dejiao Zhang, Zhiguo Wang, Andrew O. Arnold, and Bing Xiang. 2021 · 2021
Closest in time.
Increasing faithfulness in knowledge-grounded dialogue with controllable features
Rashkin, Hannah, David Reitter, Gaurav Singh Tomar, and Dipanjan Das. 2021 · 2021
Closest in time.
Santhanam, Sashank, Behnam Hedayatnia, Spandana Gella, Aishwarya Padmakumar, Seokhwan Kim, Yang Liu, and Dilek Hakkani-Tur. 2021 · 2021
Closest in time.
Evidence-based verification for real world information needs
Thorne, James, Max Glockner, Gisela Vallejo, Andreas Vlachos, and Iryna Gurevych. 2021 · 2021
Closest in time.
Faithful or extractive? on mitigating the faithfulness-abstractiveness trade-off in abstractive summarization
Ladhak, Faisal, Esin Durmus, He He, Claire Cardie, and Kathleen McKeown. 2022 · 2022
Closest in time.