Fetching the paper…
Reading the bibliography…
A common strategy for fact-checking long-form content generated by Large Language Models (LLMs) is extracting simple claims that can be verified independently.
AFaCTA: Assisting the annotation of factual claim detection with reliable LLM annotators
Jingwei Ni, Minjing Shi, Dominik Stammbach, Mrinmaya Sachan, Elliott Ash, and Markus Leippold. 2024 · 1912
Earlier work this paper cites.
NLTK: The natural language toolkit
Steven Bird and Edward Loper. 2004 · 2004
Earlier work this paper cites.
Open information extraction from the web
Michele Banko, Michael J. Cafarella, Stephen Soderland, Matt Broadhead, and Oren Etzioni. 2007 · 2007
Earlier work this paper cites.
Content Analysis: An Introduction to Its Methodology
K. Krippendorff. 2013 · 2013
Earlier work this paper cites.
Universal Decompositional Semantics on Universal Dependencies
Aaron Steven White, Drew Reisinger, Keisuke Sakaguchi, Tim Vieira, Sheng Zhang, Rachel Rudinger, Kyle Rawlins, and Benjamin Van Durme. 2016 · 2016
Earlier work this paper cites.
Fast Krippendorff: Fast computation of Krippendorff’s alpha agreement measure
Santiago Castro. 2017 · 2017
Earlier work this paper cites.
A context-aware approach for detecting worth-checking claims in political debates
Pepa Gencheva, Preslav Nakov, Lluís Màrquez, Alberto Barrón-Cedeño, and Ivan Koychev. 2017 · 2017
Earlier work this paper cites.
Toward automated fact-checking: Detecting check-worthy factual claims by claimbuster
Naeemul Hassan, Fatma Arslan, Chengkai Li, and Mark Tremayne. 2017 · 2017
Earlier work this paper cites.
An Evaluation of PredPatt and Open IE via Stage 1 Semantic Role Labeling
Sheng Zhang, Rachel Rudinger, and Ben Van Durme. 2017 · 2017
Earlier work this paper cites.
ClaimRank: Detecting check-worthy claims in Arabic and English
Israa Jaradat, Pepa Gencheva, Alberto Barrón-Cedeño, Lluís Màrquez, and Preslav Nakov. 2018 · 2018
Earlier work this paper cites.
Assessing the factual accuracy of generated text
Ben Goodrich, Vinay Rao, Peter J. Liu, and Mohammad Saleh. 2019 · 2019
Earlier work this paper cites.
A benchmark dataset of check-worthy factual claims
Fatma Arslan, Naeemul Hassan, Chengkai Li, and Mark Tremayne. 2020 · 2020
Earlier work this paper cites.
Generating fact checking briefs
Angela Fan, Aleksandra Piktus, Fabio Petroni, Guillaume Wenzek, Marzieh Saeidi, Andreas Vlachos, Antoine Bordes, and Sebastian Riedel. 2020 · 2020
Earlier work this paper cites.
Evaluating factuality in generation with dependency-level entailment
Tanya Goyal and Greg Durrett. 2020 · 2020
Earlier work this paper cites.
AmbigQA: Answering ambiguous open-domain questions
Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020 · 2020
Earlier work this paper cites.
Adversarial NLI: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020 · 2020
Earlier work this paper cites.
Fighting the COVID-19 infodemic: Modeling the perspective of journalists, fact-checkers, social media platforms, policy makers, and the society
Firoj Alam, Shaden Shaar, Fahim Dalvi, Hassan Sajjad, Alex Nikolov, Hamdy Mubarak, Giovanni Da San Martino, Ahmed Abdelali, Nadir Durrani, Kareem Darwish, Abdulaziz Al-Homaid, Wajdi Zaghouani, Tommaso Caselli, Gijs Danoe, Friso Stolk, Britt Bruntink, and Preslav Nakov. 2021 · 2021
Cited alongside, same era.
Decontextualization: Making sentences stand-alone
Eunsol Choi, Jennimaria Palomaki, Matthew Lamm, Tom Kwiatkowski, Dipanjan Das, and Michael Collins. 2021 · 2021
Cited alongside, same era.
Toward automated factchecking: Developing an annotation schema and benchmark for consistent automated claim detection
Lev Konstantinovskiy, Oliver Price, Mevan Babakar, and Arkaitz Zubiaga. 2021 · 2021
Cited alongside, same era.
Generating literal and implied subquestions to fact-check complex claims
Jifan Chen, Aniruddh Sriram, Eunsol Choi, and Greg Durrett. 2022 · 2022
Cited alongside, same era.
The clef-2022 checkthat! lab on fighting the covid-19 infodemic and fake news detection
Jacob Eisenstein, Daniel Andor, Bernd Bohnet, Michael Collins, and David Mimno. 2024 · 2024
Later among the works it cites.
AmbiFC: Fact-checking ambiguous claims with evidence
Max Glockner, Ieva Staliūnaitė, James Thorne, Gisela Vallejo, Andreas Vlachos, and Iryna Gurevych. 2024 · 2024
Later among the works it cites.
Molecular facts: Desiderata for decontextualization in LLM fact verification
Anisha Gunjal and Greg Durrett. 2024 · 2024
Later among the works it cites.
Knowledge-centric hallucination detection
Xiangkun Hu, Dongyu Ru, Lin Qiu, Qipeng Guo, Tianhang Zhang, Yang Xu, Yun Luo, Pengfei Liu, Yue Zhang, and Zheng Zhang. 2024b · 2024
Later among the works it cites.
Self-checker: Plug-and-play modules for fact-checking with large language models
Miaoran Li, Baolin Peng, Michel Galley, Jianfeng Gao, and Zhu Zhang. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Preslav Nakov, Alberto Barrón-Cedeño, Giovanni Da San Martino, Firoj Alam, Julia Maria Struß, Thomas Mandl, Rubén Míguez, Tommaso Caselli, Mucahid Kutlu, Wajdi Zaghouani, Chengkai Li, Shaden Shaar, Gautam Kishore Shahi, Hamdy Mubarak, Alex Nikolov, Nikolay Babulkov, Yavuz Selim Kartal, and Javier Beltrán. 2022 · 2022
Cited alongside, same era.
Generating scientific claims for zero-shot scientific fact checking
Dustin Wright, David Wadden, Kyle Lo, Bailey Kuehl, Arman Cohan, Isabelle Augenstein, and Lucy Lu Wang. 2022 · 2022
Cited alongside, same era.
PropSegmEnt: A large-scale corpus for proposition-level segmentation and entailment recognition
Sihao Chen, Senaka Buthpitiya, Alex Fabrikant, Dan Roth, and Tal Schuster. 2023b · 2023
Cited alongside, same era.
I-Chun Chern, Steffi Chern, Shiqi Chen, Weizhe Yuan, Kehua Feng, Chunting Zhou, Junxian He, Graham Neubig, and Pengfei Liu. 2023 · 2023
Cited alongside, same era.
WiCE: Real-world entailment for claims in Wikipedia
Ryo Kamoi, Tanya Goyal, Juan Diego Rodriguez, and Greg Durrett. 2023 · 2023
Cited alongside, same era.
We are what we repeatedly do: Inducing and deploying habitual schemas in persona-based responses
Benjamin Kane and Lenhart Schubert. 2023 · 2023
Cited alongside, same era.
Tree of clarifications: Answering ambiguous questions with retrieval-augmented large language models
Gangwoo Kim, Sungdong Kim, Byeongguk Jeon, Joonsuk Park, and Jaewoo Kang. 2023 · 2023
Cited alongside, same era.
FActScore: Fine-grained atomic evaluation of factual precision in long form text generation
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023 · 2023
Cited alongside, same era.
Claim check-worthiness detection: How well do LLMs grasp annotation guidelines?
Laura Majer and Jan Šnajder. 2024 · 2024
Later among the works it cites.
VeriScore: Evaluating the factuality of verifiable claims in long-form text generation
Yixiao Song, Yekyung Kim, and Mohit Iyyer. 2024 · 2024
Later among the works it cites.
MiniCheck: Efficient fact-checking of LLMs on grounding documents
Liyan Tang, Philippe Laban, and Greg Durrett. 2024 · 2024
Later among the works it cites.
Factcheck-bench: Fine-grained evaluation benchmark for automatic fact-checkers
Yuxia Wang, Revanth Gangi Reddy, Zain Muhammad Mujahid, Arnav Arora, Aleksandr Rubashevskii, Jiahui Geng, Osama Mohammed Afzal, Liangming Pan, Nadav Borenstein, Aditya Pillai, Isabelle Augenstein, Iryna Gurevych, and Preslav Nakov. 2024 · 2024
Later among the works it cites.
A closer look at claim decomposition
Miriam Wanner, Seth Ebner, Zhengping Jiang, Mark Dredze, and Benjamin Van Durme. 2024b · 2024
Later among the works it cites.
Long-form factuality in large language models
Jerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu, Nathan Hu, Jie Huang, Dustin Tran, Daiyi Peng, Ruibo Liu, Da Huang, Cosmo Du, and Quoc V. Le. 2024 · 2024
Later among the works it cites.
CLAMBER: A benchmark of identifying and clarifying ambiguous information needs in large language models
Tong Zhang, Peixin Qin, Yang Deng, Chen Huang, Wenqiang Lei, Junhong Liu, Dingnan Jin, Hongru Liang, and Tat-Seng Chua. 2024 · 2024
Later among the works it cites.
Factbench: A dynamic benchmark for in-the-wild language model factuality evaluation
Farima Fatahi Bayat, Lechen Zhang, Sheza Munir, and Lu Wang. 2025 · 2025
Closest in time.
A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2025 · 2025
Closest in time.
What is the essence of a claim? cross-domain claim identification
Johannes Daxenberger, Steffen Eger, Ivan Habernal, Christian Stab, and Iryna Gurevych. 2017 · 2066
Closest in time.