Fetching the paper…
Reading the bibliography…
Automatic factuality verification of large language model (LLM) generations is becoming more and more widely used to combat hallucinations.
An underspecified segmented discourse representation theory (usdrt)
Frank Schilder. 1998 · 1998
Earlier work this paper cites.
Evaluating content selection in summarization: The pyramid method
Ani Nenkova and Rebecca J Passonneau. 2004 · 2004
Earlier work this paper cites.
A corpus of general and specific sentences from news
Annie Louis and Ani Nenkova. 2012 · 2012
Earlier work this paper cites.
Improving the annotation of sentence specificity
Junyi Jessy Li, Bridget O’Daniel, Yi Wu, Wenli Zhao, and Ani Nenkova. 2016 · 2016
Earlier work this paper cites.
Ranking generated summaries by correctness: An interesting but challenging application for natural language inference
Tobias Falke, Leonardo F. R. Ribeiro, Prasetya Ajie Utama, Ido Dagan, and Iryna Gurevych. 2019 · 2019
Earlier work this paper cites.
Decontextualization: Making sentences stand-alone
Eunsol Choi, Jennimaria Palomaki, Matthew Lamm, Tom Kwiatkowski, Dipanjan Das, and Michael Collins. 2021 · 2021
Earlier work this paper cites.
Annotating and modeling fine-grained factuality in summarization
Tanya Goyal and Greg Durrett. 2021 · 2021
Earlier work this paper cites.
Measuring Attribution in Natural Language Generation Models
Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Michael Collins, Dipanjan Das, Slav Petrov, Gaurav Singh Tomar, Iulia Turc, and D. Reitter. 2021 · 2021
Earlier work this paper cites.
SituatedQA: Incorporating extra-linguistic contexts into QA
Michael J.Q. Zhang and Eunsol Choi. 2021 · 2021
Earlier work this paper cites.
Finding a Balanced Degree of Automation for Summary Evaluation
Shiyue Zhang and Mohit Bansal. 2021 · 2021
Earlier work this paper cites.
QAFactEval: Improved QA-based factual consistency evaluation for summarization
Alexander Fabbri, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2022 · 2022
Earlier work this paper cites.
Rethinking with retrieval: Faithful large language model inference
Hangfeng He, Hongming Zhang, and Dan Roth. 2022 · 2022
Earlier work this paper cites.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Andrea Madotto, and Pascale Fung. 2022 · 2022
Earlier work this paper cites.
WANLI: Worker and AI Collaboration for Natural Language Inference Dataset Creation
Alisa Liu, Swabha Swayamdipta, Noah A. Smith, and Yejin Choi. 2022 · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Cited alongside, same era.
PropSegmEnt: A Large-Scale Corpus for Proposition-level Segmentation and Entailment Recognition
Sihao Chen, Senaka Buthpitiya, Alex Fabrikant, Dan Roth, and Tal Schuster. 2023b · 2023
Cited alongside, same era.
I-Chun Chern, Steffi Chern, Shiqi Chen, Weizhe Yuan, Kehua Feng, Chunting Zhou, Junxian He, Graham Neubig, Pengfei Liu, et al. 2023 · 2023
Cited alongside, same era.
AlignScore: Evaluating Factual Consistency with A Unified Alignment Function
Yuheng Zha, Yichi Yang, Ruichen Li, and Zhiting Hu. 2023 · 2023
Later among the works it cites.
Siren’s song in the AI ocean: a survey on hallucination in large language models
Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al. 2023 · 2023
Later among the works it cites.
Complex claim verification with evidence retrieved in the wild
Jifan Chen, Grace Kim, Aniruddh Sriram, Greg Durrett, and Eunsol Choi. 2024 · 2024
Closest in time.
Merging Facts, Crafting Fallacies: Evaluating the Contradictory Nature of Aggregated Factual Claims in Long-Form Generations
Cheng-Han Chiang and Hung yi Lee. 2024 · 2024
Closest in time.
Ambigdocs: Reasoning across documents on different entities under the same name
Yoonsang Lee, Xi Ye, and Eunsol Choi. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Anisha Gunjal and Greg Durrett. 2023 · 2023
Cited alongside, same era.
WiCE: Real-World Entailment for Claims in Wikipedia
Ryo Kamoi, Tanya Goyal, Juan Rodriguez, and Greg Durrett. 2023b · 2023
Cited alongside, same era.
Evaluating Verifiability in Generative Search Engines
Nelson Liu, Tianyi Zhang, and Percy Liang. 2023a · 2023
Cited alongside, same era.
FActScore: Fine-grained atomic evaluation of factual precision in long form text generation
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023 · 2023
Cited alongside, same era.
A question answering framework for decontextualizing user-facing snippets from scientific documents
Benjamin Newman, Luca Soldaini, Raymond Fok, Arman Cohan, and Kyle Lo. 2023 · 2023
Cited alongside, same era.
Dealing with semantic underspecification in multimodal nlp
Sandro Pezzelle. 2023 · 2023
Cited alongside, same era.
Concise answers to complex questions: Summarization of long-form answers
Abhilash Potluri, Fangyuan Xu, and Eunsol Choi. 2023 · 2023
Cited alongside, same era.
Factually consistent summarization via reinforcement learning with textual entailment feedback
Paul Roit, Johan Ferret, Lior Shani, Roee Aharoni, Geoffrey Cideron, Robert Dadashi, Matthieu Geist, Sertan Girgin, Leonard Hussenot, Orgad Keller, Nikola Momchev, Sabela Ramos Garea, Piotr Stanczyk, Nino Vieillard, Olivier Bachem, Gal Elidan, Avinatan Hassidim, Olivier Pietquin, and Idan Szpektor. 2023 · 2023
Cited alongside, same era.
ExpertQA: Expert-curated questions and attributed answers
Chaitanya Malaviya, Subin Lee, Sihao Chen, Elizabeth Sieber, Mark Yatskar, and Dan Roth. 2024 · 2024
Closest in time.
MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents
Liyan Tang, Philippe Laban, and Greg Durrett. 2024 · 2024
Closest in time.
Factcheck-Bench: Fine-Grained Evaluation Benchmark for Automatic Fact-checkers
Yuxia Wang, Revanth Gangi Reddy, Zain Muhammad Mujahid, Arnav Arora, Aleksandr Rubashevskii, Jiahui Geng, Osama Mohammed Afzal, Liangming Pan, Nadav Borenstein, Aditya Pillai, Isabelle Augenstein, Iryna Gurevych, and Preslav Nakov. 2024 · 2024
Closest in time.
A Closer Look at Claim Decomposition
Miriam Wanner, Seth Ebner, Zhengping Jiang, Mark Dredze, and Benjamin Van Durme. 2024 · 2024
Closest in time.
Long-form factuality in large language models
Jerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu, Nathan Hu, Dustin Tran, Daiyi Peng, Ruibo Liu, Da Huang, Cosmo Du, et al. 2024 · 2024
Closest in time.
Fine-grained human feedback gives better rewards for language model training
Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A Smith, Mari Ostendorf, and Hannaneh Hajishirzi. 2024 · 2024
Closest in time.
How Language Model Hallucinations Can Snowball
Muru Zhang, Ofir Press, William Merrill, Alisa Liu, and Noah A. Smith. 2024 · 2024
Closest in time.