Fetching the paper…
Reading the bibliography…
Hallucinations pose a significant obstacle to the reliability and widespread adoption of language models, yet their accurate measurement remains a persistent challenge.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020 · 1919
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Alex Graves. 2012 · 2012
Earlier work this paper cites.
Multiple factor analysis by example using R
Jérôme Pagès. 2014 · 2014
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin. 2018 · 2018
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018 · 2018
Earlier work this paper cites.
A dataset for document grounded conversations
Kangyan Zhou, Shrimai Prabhumoye, and Alan W Black. 2018 · 2018
Earlier work this paper cites.
Wizard of wikipedia: Knowledge-powered conversational agents
Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2019 · 2019
Earlier work this paper cites.
Evaluating coherence in dialogue systems using entailment
Nouha Dziri, Ehsan Kamalloo, Kory Mathewson, and Osmar Zaiane. 2019 · 2019
Earlier work this paper cites.
Topical-chat: Towards knowledge-grounded open-domain conversations
Karthik Gopalakrishnan, Behnam Hedayatnia, Qinlang Chen, Anna Gottardi, Sanjeev Kwatra, Anu Venkatesh, Raefer Gabriel, and Dilek Hakkani-Tür. 2019 · 2019
Earlier work this paper cites.
OpenDialKG: Explainable conversational reasoning with attention-based walks over knowledge graphs
Seungwhan Moon, Pararth Shah, Anuj Kumar, and Rajen Subba. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization
Esin Durmus, He He, and Mona Diab. 2020 · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Earlier work this paper cites.
Asking and answering questions to evaluate the factual consistency of summaries
Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020 · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Earlier work this paper cites.
Neural path hunter: Reducing hallucination in dialogue systems via path grounding
Nouha Dziri, Andrea Madotto, Osmar Zaïane, and Avishek Joey Bose. 2021 · 2021
Earlier work this paper cites.
q 2 q^{2} : Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering
Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend. 2021 · 2021
Earlier work this paper cites.
Focused attention improves document-grounded generation
Shrimai Prabhumoye, Kazuma Hashimoto, Yingbo Zhou, Alan W Black, and Ruslan Salakhutdinov. 2021 · 2021
Earlier work this paper cites.
Increasing faithfulness in knowledge-grounded dialogue with controllable features
Hannah Rashkin, David Reitter, Gaurav Singh Tomar, and Dipanjan Das. 2021 · 2021
Earlier work this paper cites.
The curious case of hallucinations in neural machine translation
Vikas Raunak, Arul Menezes, and Marcin Junczys-Dowmunt. 2021 · 2021
Earlier work this paper cites.
Hallucinated but factual! inspecting the factuality of hallucinations in abstractive summarization
Meng Cao, Yue Dong, and Jackie Cheung. 2022 · 2022
Cited alongside, same era.
Diving deep into modes of fact hallucinations in dialogue systems
Souvik Das, Sougata Saha, and Rohini K Srihari. 2022 · 2022
Cited alongside, same era.
Spurious correlations in reference-free evaluation of text generation
Esin Durmus, Faisal Ladhak, and Tatsunori Hashimoto. 2022 · 2022
Cited alongside, same era.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre. 2022 · 2022
Cited alongside, same era.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Llms instead of human judges? a large scale empirical study across 20 nlp evaluation tasks
Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, André F. T. Martins, Philipp Mondorf, Vera Neplenbroek, Sandro Pezzelle, Barbara Plank, David Schlangen, Alessandro Suglia, Aditya K Surikuchi, Ece Takmaz, and Alberto Testoni. 2024 · 2024
Later among the works it cites.
Dola: Decoding by contrasting layers improves factuality in large language models
Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James R. Glass, and Pengcheng He. 2024 · 2024
Later among the works it cites.
Haloscope: Harnessing unlabeled llm generations for hallucination detection
Xuefeng Du, Chaowei Xiao, and Yixuan Li. 2024 · 2024
Later among the works it cites.
Shifting attention to relevance: Towards the predictive uncertainty quantification of free-form large language models
Jinhao Duan, Hao Cheng, Shiqi Wang, Alex Zavalny, Chenan Wang, Renjing Xu, Bhavya Kailkhura, and Kaidi Xu. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022 · 2022
Cited alongside, same era.
Broken neural scaling laws
Ethan Caballero, Kshitij Gupta, Irina Rish, and David Krueger. 2023 · 2023
Cited alongside, same era.
Hallucination detection: Robustly discerning reliable answers in large language models
Yuyan Chen, Qiang Fu, Yichen Yuan, Zhihao Wen, Ge Fan, Dayiheng Liu, Dongmei Zhang, Zhixu Li, and Yanghua Xiao. 2023 · 2023
Cited alongside, same era.
Chain-of-verification reduces hallucination in large language models
Shehzaad Dhuliawala, Mojtaba Komeili, Jing Xu, Roberta Raileanu, Xian Li, Asli Celikyilmaz, and Jason Weston. 2023 · 2023
Cited alongside, same era.
Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Harkirat Singh Behl, Xin Wang, Sébastien Bubeck, Ronen Eldan, Adam Tauman Kalai, Yin Tat Lee, and Yuanzhi Li. 2023 · 2023
Cited alongside, same era.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023 · 2023
Cited alongside, same era.
Abhimanyu Dubey et al. 2024 · 2024
Later among the works it cites.
Gemma: Open models based on gemini research and technology
Team Gemma. 2024 · 2024
Later among the works it cites.
Understanding finetuning for factual knowledge extraction
Gaurav Rohit Ghosal, Tatsunori Hashimoto, and Aditi Raghunathan. 2024 · 2024
Later among the works it cites.
OLMo: Accelerating the science of language models
Dirk Groeneveld et al. 2024 · 2024
Later among the works it cites.
Calibrated language models must hallucinate
Adam Tauman Kalai and Santosh S. Vempala. 2024 · 2024
Later among the works it cites.
Comparing hallucination detection metrics for multilingual generation
Haoqiang Kang, Terra Blevins, and Luke Zettlemoyer. 2024 · 2024
Later among the works it cites.
The dawn after the dark: An empirical study on factuality hallucination in large language models
Junyi Li, Jie Chen, Ruiyang Ren, Xiaoxue Cheng, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. 2024 · 2024
Later among the works it cites.
Exploring and evaluating hallucinations in llm-powered code generation
Fang Liu, Yang Liu, Lin Shi, Houkun Huang, Ruifeng Wang, Zhen Yang, Li Zhang, Zhongqi Li, and Yuchi Ma. 2024 · 2024
Later among the works it cites.
Hallucination detection and hallucination mitigation: An investigation
Junliang Luo, Tianyu Li, Di Wu, Michael Jenkin, Steve Liu, and Gregory Dudek. 2024 · 2024
Later among the works it cites.
Team OpenAI. 2024 · 2024
Later among the works it cites.
Trusting your evidence: Hallucinate less with context-aware decoding
Weijia Shi, Xiaochuang Han, Mike Lewis, Yulia Tsvetkov, Luke Zettlemoyer, and Wen-tau Yih. 2024 · 2024
Later among the works it cites.
Sayself: Teaching llms to express confidence with self-reflective rationales
Tianyang Xu, Shujin Wu, Shizhe Diao, Xiaoze Liu, Xingyao Wang, Yangyi Chen, and Jing Gao. 2024a · 2024
Later among the works it cites.
Evaluating the effectiveness of llm-evaluators (aka llm-as-judge)
Ziyou Yan. 2024 · 2024
Later among the works it cites.
Verify with caution: The pitfalls of relying on imperfect factuality metrics
Ameya Godbole and Robin Jia. 2025 · 2025
Closest in time.
A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2025 · 2025
Closest in time.
Judging the judges: Evaluating alignment and vulnerabilities in llms-as-judges
Aman Singh Thakur, Kartik Choudhary, Venkat Srinik Ramayapally, Sankaran Vaidyanathan, and Dieuwke Hupkes. 2025 · 2025
Closest in time.
Saad Obaid ul Islam, Anne Lauscher, and Goran Glavaš. 2025 · 2025
Closest in time.
Hallucinate at the last in long response generation: A case study on long document summarization
Joonho Yang, Seunghyun Yoon, Hwan Chang, Byeongjeong Kim, and Hwanhee Lee. 2025 · 2025
Closest in time.
Towards a unified multi-dimensional evaluator for text generation
Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022 · 2038
Closest in time.