Fetching the paper…
Reading the bibliography…
The rapid advancement of Large Language Models (LLMs) has brought a pressing challenge: how to reliably assess hallucinations to guarantee model trustworthiness.
Embedding and Gradient Say Wrong: A White-box Method for Hallucination Detection
Xiaomeng Hu, Yiming Zhang, Ru Peng, Haozhe Zhang, Chenwei Wu, Gang Chen, and Junbo Zhao · 1959
Earlier work this paper cites.
Climate-fever: A dataset for verification of real-world climate claims
Thomas Diggelmann, Jordan Boyd-Graber, Jannis Bulian, Massimiliano Ciaramita, and Markus Leippold · 2012
Earlier work this paper cites.
Problems in current text simplification research: New data can help
Wei Xu, Chris Callison-Burch, and Courtney Napoles · 2015
Earlier work this paper cites.
Get to the point: Summarization with pointer-generator networks
Abigail See, Peter J. Liu, and Christopher D. Manning · 2017
Earlier work this paper cites.
Sentence simplification with deep reinforcement learning
Xingxing Zhang and Mirella Lapata · 2017
Earlier work this paper cites.
Survey of the State of the Art in Natural Language Generation: Core tasks, applications and evaluation
Albert Gatt and Emiel Krahmer · 2018
Earlier work this paper cites.
Hallucinations in neural machine translation
Katherine Lee, Orhan Firat, Ashish Agarwal, Clara Fannjiang, and David Sussillo · 2018
Earlier work this paper cites.
Bootstrapping generators from noisy data
Laura Perez-Beltrachini and Mirella Lapata · 2018
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Handling divergent reference texts when evaluating table-to-text generation
Bhuwan Dhingra, Manaal Faruqui, Ankur Parikh, Ming-Wei Chang, Dipanjan Das, and William W Cohen · 2019
Earlier work this paper cites.
Wizard of wikipedia: Knowledge-powered conversational agents
Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston · 2019
Earlier work this paper cites.
Ranking generated summaries by correctness: An interesting but challenging application for natural language inference
Tobias Falke, Leonardo F. R. Ribeiro, Prasetya Ajie Utama, Ido Dagan, and Iryna Gurevych · 2019
Earlier work this paper cites.
Assessing the factual accuracy of generated text
Ben Goodrich, Vinay Rao, Peter J. Liu, and Mohammad Saleh · 2019
Earlier work this paper cites.
Topical-chat: Towards knowledge-grounded open-domain conversations
Karthik Gopalakrishnan, Behnam Hedayatnia, Qinglang Chen, Anna Gottardi, Sanjeev Kwatra, Anu Venkatesh, Raefer Gabriel, and Dilek Hakkani-Tür · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel · 2019
Earlier work this paper cites.
Best practices for the human evaluation of automatically generated text
Chris van der Lee, Albert Gatt, Emiel van Miltenburg, Sander Wubben, and Emiel Krahmer · 2019
Earlier work this paper cites.
Dialogue natural language inference
Sean Welleck, Jason Weston, Arthur Szlam, and Kyunghyun Cho · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization
Esin Durmus, He He, and Mona Diab · 2020
Earlier work this paper cites.
Evaluating semantic accuracy of data-to-text generation with natural language inference
Ondřej Dušek and Zdeněk Kasner · 2020
Earlier work this paper cites.
Evaluating factuality in generation with dependency-level entailment
Tanya Goyal and Greg Durrett · 2020
Earlier work this paper cites.
Twenty years of confusion in human evaluation: NLG needs evaluation sheets and standardised definitions
David M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, and Verena Rieser · 2020
Earlier work this paper cites.
What have we achieved on text summarization?
Dandan Huang, Leyang Cui, Sen Yang, Guangsheng Bao, Kun Wang, Jun Xie, and Yue Zhang · 2020
Earlier work this paper cites.
Evaluating the factual consistency of abstractive text summarization
Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher · 2020
Earlier work this paper cites.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald · 2020
Earlier work this paper cites.
Asking and answering questions to evaluate the factual consistency of summaries
Alex Wang, Kyunghyun Cho, and Mike Lewis · 2020
Earlier work this paper cites.
Are factuality checkers reliable? adversarial meta-evaluation of factuality in summarization
Yiran Chen, Pengfei Liu, and Xipeng Qiu · 2021
Earlier work this paper cites.
SummEval: Re-evaluating summarization evaluation
Alexander R. Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev · 2021
Earlier work this paper cites.
GO FIGURE: A meta evaluation of factuality in summarization
Saadia Gabriel, Asli Celikyilmaz, Rahul Jha, Yejin Choi, and Jianfeng Gao · 2021
Earlier work this paper cites.
q 2 q^{2} : Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering
Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend · 2021
Earlier work this paper cites.
The Factual Inconsistency Problem in Abstractive Text Summarization: A Survey
Yi-Chong Huang, Xia-Chong Feng, Xiao-Cheng Feng, and Bing Qin · 2021
Earlier work this paper cites.
Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics
Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov · 2021
Earlier work this paper cites.
Don’t be contradicted with anything! CI-ToD: Towards benchmarking consistency for task-oriented dialogue system
Libo Qin, Tianbao Xie, Shijue Huang, Qiguang Chen, Xiao Xu, and Wanxiang Che · 2021
Earlier work this paper cites.
The curious case of hallucinations in neural machine translation
Vikas Raunak, Arul Menezes, and Marcin Junczys-Dowmunt · 2021
Earlier work this paper cites.
Controlling the factual correctness of a data-to-text generator by adding a plain-text description
Clément Rebuffel, Thomas Robert, Geoffrey Bernard, and Laure Soulier · 2021
Earlier work this paper cites.
BEAMetrics: A Benchmark for Language Generation Evaluation Evaluation
Thomas Scialom and Felix Hill · 2021
Earlier work this paper cites.
QuestEval: Summarization asks for fact-based evaluation
Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari · 2021
Earlier work this paper cites.
Generation challenges: Results of the accuracy evaluation shared task
Craig Thomson and Ehud Reiter · 2021
Earlier work this paper cites.
Factual consistency evaluation for text summarization via counterfactual estimation
Yuexiang Xie, Fei Sun, Yang Deng, Yaliang Li, and Bolin Ding · 2021
Earlier work this paper cites.
Hallucinated but factual! inspecting the factuality of hallucinations in abstractive summarization
Meng Cao, Yue Dong, and Jackie Cheung · 2022
Earlier work this paper cites.
Evaluating factuality in text simplification
Ashwin Devaraj, William Sheffield, Byron Wallace, and Junyi Jessy Li · 2022
Earlier work this paper cites.
Faithful to the document or to the world? mitigating hallucinations via entity-linked knowledge in abstractive summarization
Yue Dong, John Wieting, and Pat Verga · 2022
Earlier work this paper cites.
Is GPT-3 text indistinguishable from human text? scarecrow: A framework for scrutinizing machine text
Yao Dou, Maxwell Forbes, Rik Koncel-Kedziorski, Noah A. Smith, and Yejin Choi · 2022
Earlier work this paper cites.
FaithDial: A faithful benchmark for information-seeking dialogue
Nouha Dziri, Ehsan Kamalloo, Sivan Milton, Osmar Zaiane, Mo Yu, Edoardo M. Ponti, and Siva Reddy · 2022
Earlier work this paper cites.
Evaluating attribution in dialogue systems: The BEGIN benchmark
Nouha Dziri, Hannah Rashkin, Tal Linzen, and David Reitter · 2022
Earlier work this paper cites.
QAFactEval: Improved QA-based factual consistency evaluation for summarization
Alexander Fabbri, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong · 2022
Earlier work this paper cites.
DialSummEval: Revisiting summarization evaluation for dialogues
Mingqi Gao and Xiaojun Wan · 2022
Earlier work this paper cites.
PRISMA2020: An R package and Shiny app for producing PRISMA 2020-compliant flow diagrams, with interactivity for optimised digital transparency and Open Synthesis
Neal R. Haddaway, Matthew J. Page, C. C. Pritchard, and Luke A. McGuinness · 2022
Earlier work this paper cites.
TRUE: Re-evaluating factual consistency evaluation
Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansy, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, and Yossi Matias · 2022
Earlier work this paper cites.
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Guzman, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jared Kaplan, Catherine Olsson Janus, Sam Ringer, Shaun E. Tan, Jared M. Sugarman, Benjamin A. G. Graham, Neel Nanda, Kamal Ndousse, William Saunders, Jackson Kernion, Liane Lovitt, Christopher Olah, Ben Mann, Dario Amodei, Tom B. Brown, Jack Clark, Nicholas Joseph, Samuel R. McCandlish, Chris Olah, and Sam McCandlish · 2022
Earlier work this paper cites.
SummaC: Re-visiting NLI-based models for inconsistency detection in summarization
Philippe Laban, Tobias Schnabel, Paul N. Bennett, and Marti A. Hearst · 2022
Earlier work this paper cites.
Masked summarization to generate factually inconsistent summaries for improved factual consistency checking
Hwanhee Lee, Kang Min Yoo, Joonsuk Park, Hwaran Lee, and Kyomin Jung · 2022
Earlier work this paper cites.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Earlier work this paper cites.
A token-level reference-free hallucination detection benchmark for free-form text generation
Tianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao, Zhifang Sui, Weizhu Chen, and Bill Dolan · 2022
Earlier work this paper cites.
Chatgpt, 2022
OpenAI · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe · 2022
Earlier work this paper cites.
FactGraph: Evaluating factuality in summarization with semantic graph representations
Leonardo F. R. Ribeiro, Mengwen Liu, Iryna Gurevych, Markus Dreyer, and Mohit Bansal · 2022
Earlier work this paper cites.
Harim + {}^{\mbox{+}} : Evaluating summary quality with hallucination risk
Seonil Son, Junsoo Park, Jeong-In Hwang, Junghwa Lee, Hyungjong Noh, and Yeonsoo Lee · 2022
Earlier work this paper cites.
Falsesum: Generating document-level NLI examples for recognizing factual inconsistency in summarization
Prasetya Utama, Joshua Bambrick, Nafise Moosavi, and Iryna Gurevych · 2022
Earlier work this paper cites.
Analyzing and evaluating faithfulness in dialogue summarization
Bin Wang, Chen Zhang, Yan Zhang, Yiming Chen, and Haizhou Li · 2022
Earlier work this paper cites.
Shared Imagination: LLMs Hallucinate Alike
Yilun Zhou, Caiming Xiong, Silvio Savarese, and Chien-Sheng Wu · 2022
Earlier work this paper cites.
A meta-evaluation of faithfulness metrics for long-form hospital-course summarization
Griffin Adams, Jason Zuckerg, and Noémie Elhadad · 2023
Cited alongside, same era.
Autohall: Automated hallucination dataset generation for large language models
Zouying Cao, Yifei Yang, and Hai Zhao · 2023
Cited alongside, same era.
Beyond factuality: A comprehensive evaluation of large language models as knowledge generators
Liang Chen, Yang Deng, Yatao Bian, Zeyu Qin, Bingzhe Wu, Tat-Seng Chua, and Kam-Fai Wong · 2023
Cited alongside, same era.
FELM: benchmarking factuality evaluation of large language models
Shiqi Chen, Yiran Zhao, Jinghan Zhang, I-Chun Chern, Siyang Gao, Pengfei Liu, and Junxian He · 2023
Cited alongside, same era.
Evaluating Hallucinations in Chinese Large Language Models
Qinyuan Cheng, Tianxiang Sun, Wenwei Zhang, Siyin Wang, Xiangyang Liu, Mozhi Zhang, Junliang He, Mianqiu Huang, Zhangyue Yin, Kai Chen, and Xipeng Qiu · 2023
Cited alongside, same era.
Comparing Hallucination Detection Metrics for Multilingual Generation
Haoqiang Kang, Terra Blevins, and Luke Zettlemoyer · 2024
Closest in time.
From general to specific: Utilizing general hallucination to benchmark specific role-playing agents
Chuyi Kong, Ziyang Luo, Hongzhan Lin, Zhiyuan Fan, Yaxin Fan, Yuxi Sun, and Jing Ma · 2024
Closest in time.
Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
Jannik Kossen, Jiatong Han, Muhammed Razzak, Lisa Schut, Shreshth A. Malik, and Yarin Gal · 2024
Closest in time.
Materials science in the era of large language models: a perspective
Ge Lei, Ronan Docherty, and Samuel J Cooper · 2024
Closest in time.
UHGEval: Benchmarking the Hallucination of Chinese Large Language Models via Unconstrained Generation
Xun Liang, Shichao Song, Simin Niu, Zhiyu Li, Feiyu Xiong, Bo Tang, Yezhaohui Wang, Dawei He, Cheng Peng, Zhonghao Wang, and Haiying Deng · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
FacTool: Factuality Detection in Generative AI - A Tool Augmented Framework for Multi-task and Multi-domain Scenarios
I-Chun Chern, Steffi Chern, Shiqi Chen, Weizhe Yuan, Kehua Feng, Chunting Zhou, Junxian He, Graham Neubig, and Pengfei Liu · 2023
Cited alongside, same era.
LM vs LM: Detecting factual errors via cross examination
Roi Cohen, May Hamri, Mor Geva, and Amir Globerson · 2023
Cited alongside, same era.
Chatlaw: Open-source legal large language model with integrated external knowledge bases
Jiaxi Cui, Zongjian Li, Yang Yan, Bohua Chen, and Li Yuan · 2023
Cited alongside, same era.
Detecting and mitigating hallucinations in machine translation: Model internal workings alone do well, sentence similarity Even better
David Dale, Elena Voita, Loic Barrault, and Marta R. Costa-jussà · 2023
Cited alongside, same era.
HalOmi: A manually annotated benchmark for multilingual hallucination and omission detection in machine translation
David Dale, Elena Voita, Janice Lam, Prangthip Hansanti, Christophe Ropers, Elahe Kalbassi, Cynthia Gao, Loic Barrault, and Marta Costa-jussà · 2023
Cited alongside, same era.
FactKB: Generalizable factuality evaluation using language models enhanced with factual knowledge
Shangbin Feng, Vidhisha Balachandran, Yuyang Bai, and Yulia Tsvetkov · 2023
Cited alongside, same era.
Chainpoll: A high efficacy method for LLM hallucination detection
Robert Friel and Atindriyo Sanyal · 2023
Cited alongside, same era.
Exploring and evaluating hallucinations in llm-powered code generation
Fang Liu, Yang Liu, Lin Shi, Houkun Huang, Ruifeng Wang, Zhen Yang, and Li Zhang · 2024
Closest in time.
SummaCoz: A Dataset for Improving the Interpretability of Factual Consistency Detection for Summarization
Ge Luo, Weisi Fan, Miaoran Li, Guoruizhe Sun, Runlong Zhang, Chenyu Xu, and Forrest Sheng Bao · 2024
Closest in time.
Metacheckgpt - A multi-task hallucination detector using LLM uncertainty and meta-models
Rahul Mehta, Andrew Hoblitzell, Jack O’Keefe, Hyeju Jang, and Vasudeva Varma · 2024
Closest in time.
Generating benchmarks for factuality evaluation of language models
Dor Muhlgay, Ori Ram, Inbal Magar, Yoav Levine, Nir Ratner, Yonatan Belinkov, Omri Abend, Kevin Leyton-Brown, Amnon Shashua, and Yoav Shoham · 2024
Closest in time.
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-augmented Language Models
Cheng Niu, Yuanhao Wu, Juno Zhu, Siliang Xu, KaShun Shum, Randy Zhong, Juntong Song, and Tong Zhang · 2024
Closest in time.
Erbench: An entity-relationship based automatically verifiable hallucination benchmark for large language models
Jio Oh, Soyeon Kim, Junseok Seo, Jindong Wang, Ruochen Xu, Xing Xie, and Steven Whang · 2024
Closest in time.
Fine-tuning or Retrieval? Comparing Knowledge Injection in LLMs
Oded Ovadia, Menachem Brief, Moshik Mishaeli, and Oren Elisha · 2024
Closest in time.
Defan: Definitive answer dataset for llms hallucination evaluation
A B. M. Ashikur Rahman, Saeed Anwar, Muhammad Usman, and Ajmal Mian · 2024
Closest in time.
MASSIVE multilingual Abstract Meaning Representation: A dataset and baselines for hallucination detection
Michael Regan, Shira Wein, George Baker, and Emilio Monti · 2024
Closest in time.
The Human Factor in Detecting Errors of Large Language Models: A Systematic Literature Review and Future Research Directions
Christian A. Schiller · 2024
Closest in time.
FENICE: factuality evaluation of summarization based on natural language inference and claim extraction
Alessandro Scirè, Karim Ghonim, and Roberto Navigli · 2024
Closest in time.
Jean Seo, Jongwon Lim, Dongjun Jang, and Hyopil Shin · 2024
Closest in time.
AXCEL: Automated eXplainable Consistency Evaluation using LLMs
P. Aditya Sreekar, Sahil Verma, Suransh Chopra, Abhishek Persad, Sarik Ghazarian, and Narayanan Sadagopan · 2024
Closest in time.
Llm-check: Investigating detection of hallucinations in large language models
Gaurang Sriramanan, Siddhant Bharti, Vinu Sankar Sadasivan, Shoumik Saha, Priyatham Kattakinda, and Soheil Feizi · 2024
Closest in time.
Unsupervised Real-time Hallucination Detection based on the Internal States of Large Language Models
Weihang Su, Changyue Wang, Qingyao Ai, Yiran Hu, Zhijing Wu, Yujia Zhou, and Yiqun Liu · 2024
Closest in time.
Benchmarking hallucination in large language models based on unanswerable math word problem
YuHong Sun, Zhangyue Yin, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Hui Zhao · 2024
Closest in time.
TofuEval: Evaluating hallucinations of LLMs on topic-focused dialogue summarization
Liyan Tang, Igor Shalyminov, Amy Wong, Jon Burnsky, Jake Vincent, Yu’an Yang, Siffi Singh, Song Feng, Hwanjun Song, Hang Su, Lijia Sun, Yi Zhang, Saab Mansour, and Kathleen McKeown · 2024
Closest in time.
FreshLLMs: Refreshing Large Language Models with Search Engine Augmentation
Tu Vu, Mohit Iyyer, Xuezhi Wang, Noah Constant, Jerry W. Wei, Jason Wei, Chris Tar, Yun-Hsuan Sung, Denny Zhou, Quoc V. Le, and Thang Luong · 2024
Closest in time.
Acueval: Fine-grained hallucination evaluation and correction for abstractive summarization
David Wan, Koustuv Sinha, Srini Iyer, Asli Celikyilmaz, Mohit Bansal, and Ramakanth Pasunuru · 2024
Closest in time.
Long-form factuality in large language models
Jerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu, Nathan Hu, Jie Huang, Dustin Tran, Daiyi Peng, Ruibo Liu, Da Huang, Cosmo Du, and Quoc V. Le · 2024
Closest in time.
Zikai Xie · 2024
Closest in time.
InterrogateLLM: Zero-resource Hallucination Detection in LLM-generated Answers
Yakir Yehuda, Itzik Malkiel, Oren Barkan, Jonathan Weill, Royi Ronen, and Noam Koenigstein · 2024
Closest in time.
Kola: Carefully benchmarking world knowledge of large language models
Jifan Yu, Xiaozhi Wang, Shangqing Tu, Shulin Cao, Daniel Zhang-Li, Xin Lv, Hao Peng, Zijun Yao, Xiaohan Zhang, Hanming Li, Chunyang Li, Zheyuan Zhang, Yushi Bai, Yantao Liu, Amy Xin, Kaifeng Yun, Linlu Gong, Nianyi Lin, Jianhui Chen, Zhili Wu, Yunjia Qi, Weikai Li, Yong Guan, Kaisheng Zeng, Ji Qi, Hailong Jin, Jinxin Liu, Yu Gu, Yuan Yao, Ning Ding, Lei Hou, Zhiyuan Liu, Bin Xu, Jie Tang, and Juanzi Li · 2024
Closest in time.
ReEval: Automatic hallucination evaluation for retrieval-augmented large language models via transferable adversarial attacks
Xiaodong Yu, Hao Cheng, Xiaodong Liu, Dan Roth, and Jianfeng Gao · 2024
Closest in time.
Whispers that Shake Foundations: Analyzing and Mitigating False Premise Hallucinations in Large Language Models
Hongbang Yuan, Pengfei Cao, Zhuoran Jin, Yubo Chen, Daojian Zeng, Kang Liu, and Jun Zhao · 2024
Closest in time.
Fine-grained natural language inference based faithfulness evaluation for diverse summarisation tasks
Huajian Zhang, Yumo Xu, and Laura Perez-Beltrachini · 2024
Closest in time.
Fine-grained natural language inference based faithfulness evaluation for diverse summarisation tasks
Huajian Zhang, Yumo Xu, and Laura Perez-Beltrachini · 2024
Closest in time.
Leveraging entailment judgements in cross-lingual summarisation
Huajian Zhang, Yumo Xu, and Laura Perez-Beltrachini · 2024
Closest in time.
Halluverse25: Fine-grained multilingual benchmark dataset for LLM hallucinations
Samir Abdaljalil, Hasan Kurban, and Erchin Serpedin · 2025
Closest in time.
HalluLens: LLM Hallucination Benchmark
Yejin Bang, Ziwei Ji, Alan Schelten, Anthony Hartshorn, Tara Fowler, Cheng Zhang, Nicola Cancedda, and Pascale Fung · 2025
Closest in time.
Factbench: A dynamic benchmark for in-the-wild language model factuality evaluation
Farima Fatahi Bayat, Lechen Zhang, Sheza Munir, and Lu Wang · 2025
Closest in time.
Luna: A lightweight evaluation model to catch language model hallucinations with high accuracy and low cost
Masha Belyi, Robert Friel, Shuai Shao, and Atindriyo Sanyal · 2025
Closest in time.
Jiahao Cheng, Tiancheng Su, Jia Yuan, Guoxiu He, Jiawei Liu, Xinqi Tao, Jingwen Xie, and Huaxia Li · 2025
Closest in time.
Evaluation hallucination in multi-round incomplete information lateral-driven reasoning tasks
Wenhan Dong, Tianyi Hu, Jingyi Zheng, Zhen Sun, Yuemeng Zhao, Yule Liu, Xinlei He, and Xinyi Huang · 2025
Closest in time.
Hallumix: A task-agnostic, multi-domain benchmark for real-world hallucination detection
Deanna Emery, Michael Goitia, Freddie Vargus, and Iulia Neagu · 2025
Closest in time.
Uncertainty quantification in retrieval augmented question answering
Shashank Gupta, Devamanyu Hazarika, Gaurav Sikka, Parth Sarthi, Rogerio Feris, Kate Saenko, Sathya N. Ravi, and Vikas Singh · 2025
Closest in time.
Beyond facts: Evaluating intent hallucination in large language models
Yijie Hao, Haofei Yu, and Jiaxuan You · 2025
Closest in time.
Chinese simpleqa: A chinese factuality evaluation for large language models
Yancheng He, Shilong Li, Jiaheng Liu, Yingshui Tan, Weixun Wang, Hui Huang, Xingyuan Bu, Hangyu Guo, Chengwei Hu, Boren Zheng, Zhuoran Lin, Dekai Sun, Zhicheng Zheng, Wenbo Su, and Bo Zheng · 2025
Closest in time.
Medscore: Factuality evaluation of free-form medical answers
Heyuan Huang, Alexandra DeLucia, Vijay Murari Tiyyala, and Mark Dredze · 2025
Closest in time.
Bi’an: A bilingual benchmark and model for hallucination detection in retrieval-augmented generation
Zhouyu Jiang, Mengshu Sun, Zhiqiang Zhang, and Lei Liang · 2025
Closest in time.
On A scale from 1 to 5: Quantifying hallucination in faithfulness evaluation
Xiaonan Jing, Srinivas Billa, and Danny Godbout · 2025
Closest in time.
Evaluating evaluation metrics - the mirage of hallucination detection
Atharva Kulkarni, Yuan Zhang, Joel Ruben Antony Moniz, Xiou Ge, Bo-Hsiang Tseng, Dhivya Piraviperumal, Swabha Swayamdipta, and Hong Yu · 2025
Closest in time.
Multihal: Multilingual dataset for knowledge-graph grounded evaluation of LLM hallucinations
Ernests Lavrinovics, Russa Biswas, Katja Hose, and Johannes Bjerva · 2025
Closest in time.
How llms react to industrial spatio-temporal data? assessing hallucination with a novel traffic incident benchmark dataset
Qiang Li, Mingkun Tan, Xun Zhao, Dan Zhang, Daoan Zhang, Shengzhao Lei, Anderson S. Chu, Lujun Li, and Porawit Kamnoedboon · 2025
Closest in time.
Treecut: A synthetic unanswerable math word problem dataset for LLM hallucination evaluation
Jialin Ouyang · 2025
Closest in time.
MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models
Shrey Pandit, Jiawei Xu, Junyuan Hong, Zhangyang Wang, Tianlong Chen, Kaidi Xu, and Ying Ding · 2025
Closest in time.
HalluciNot: Hallucination Detection Through Context and Common Knowledge Verification, 2025
Bibek Paudel, Alexander Lyzhov, Preetam Joshi, and Puneet Anand · 2025
Closest in time.
Evaluating LLMs’ Assessment of Mixed-context Hallucination Through the Lens of Summarization
Siya Qi, Rui Cao, Yulan He, and Zheng Yuan · 2025
Closest in time.
Verifastscore: Speeding up long-form factuality evaluation
Rishanth Rajendhran, Amir Zadeh, Matthew Sarte, Chuan Li, and Mohit Iyyer · 2025
Closest in time.
K-HALU: multiple answer korean hallucination benchmark for large language models
Jaehyung Seo and Heuiseok Lim · 2025
Closest in time.
T2F: a domain-agnostic multi-agent framework for unstructured text to factuality evaluation items generation
Xin Tong, Jingya Wang, Yasen Aizezi, Hanming Zhai, and Bo Jin · 2025
Closest in time.
How Much Do LLMs Hallucinate across Languages? On Multilingual Estimation of LLM Hallucination in the Wild
Saad Obaid ul Islam, Anne Lauscher, and Goran Glavas · 2025
Closest in time.
Changyue Wang, Weihang Su, Qingyao Ai, and Yiqun Liu · 2025
Closest in time.
Guangzhi Xiong, Eric Xie, Corey Williams, Myles Kim, Amir Hassan Shariatmadari, Sikun Guo, Stefan Bekiranov, and Aidong Zhang · 2025
Closest in time.
Zhiwen You and Yue Guo · 2025
Closest in time.
HaDeMiF: Hallucination Detection and Mitigation in Large Language Models
Xiaoling Zhou, Mingjie Zhang, Zhemg Lee, Wei Ye, and Shikun Zhang · 2025
Closest in time.