Fetching the paper…
Reading the bibliography…
Large language models (LLMs) often produce errors, including factual inaccuracies, biases, and reasoning failures, collectively referred to as "hallucinations".
Data mining for detecting errors in dictation speech recognition
Lina Zhou, Yongmei Shi, Jinjuan Feng, and Andrew Sears · 2005
Earlier work this paper cites.
Error detection in confusion network
Alexandre Allauzen · 2007
Earlier work this paper cites.
Error detection in broadcast news ASR using markov chains
Thomas Pellegrini and Isabel Trancoso · 2009
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
ASR error detection in a conversational spoken language translation system
Wei Chen, Sankaranarayanan Ananthakrishnan, Rohit Kumar, Rohit Prasad, and Prem Natarajan · 2013
Earlier work this paper cites.
A survey of spelling error detection and correction techniques
Ritika Mishra and Navjot Kaur · 2013
Earlier work this paper cites.
Uw-stanford system description for AESW 2016 shared task on grammatical error detection
Dan Flickinger, Michael Wayne Goodman, and Woodley Packard · 2016
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
Automatic speech recognition errors detection and correction: A review
Rahhal Errattahi, Asmaa El Hannani, and Hassan Ouahmane · 2018
Earlier work this paper cites.
Wronging a right: Generating better errors to improve grammatical error detection
Sudhanshu Kasewa, Pontus Stenetorp, and Sebastian Riedel · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D Manning · 2018
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang · 2018
Earlier work this paper cites.
Context is key: Grammatical error detection with contextual word representations
Samuel J. Bell, Helen Yannakoudakis, and Marek Rei · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Matthew Kelcey, Jacob Devlin, Kenton Lee, Kristina N. Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov · 2019
Earlier work this paper cites.
YiSi - a unified semantic MT quality evaluation and estimation metric for languages with different levels of available resources
Chi-kiu Lo · 2019
Earlier work this paper cites.
ERNIE: Enhanced language representation with informative entities
Zhengyan Zhang, Xu Han, Zhiyuan Liu, Xin Jiang, Maosong Sun, and Qun Liu · 2019
Earlier work this paper cites.
On identifiability in transformers
Gino Brunner, Yang Liu, Damian Pascual, Oliver Richter, Massimiliano Ciaramita, and Roger Wattenhofer · 2020
Earlier work this paper cites.
Grammatical error detection in transcriptions of spoken english
Andrew Caines, Christian Bentz, Kate M. Knill, Marek Rei, and Paula Buttery · 2020
Earlier work this paper cites.
Chinese grammatical error detection based on BERT model
Yong Cheng and Mofan Duan · 2020
Earlier work this paper cites.
KoBE: Knowledge-based machine translation evaluation
Zorik Gekhman, Roee Aharoni, Genady Beryozkin, Markus Freitag, and Wolfgang Macherey · 2020
Earlier work this paper cites.
Evaluating the factual consistency of abstractive text summarization
Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
Overview of NLPTEA-2020 shared task for Chinese grammatical error diagnosis
Gaoqi Rao, Erhong Yang, and Baolin Zhang · 2020
Earlier work this paper cites.
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie · 2020
Earlier work this paper cites.
Anthropomorphism in ai
Arleen Salles, Kathinka Evers, and Michele Farisco · 2020
Earlier work this paper cites.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh · 2020
Earlier work this paper cites.
On exposure bias, hallucination and domain shift in neural machine translation
Chaojun Wang and Rico Sennrich · 2020
Earlier work this paper cites.
Grammatical error detection with self attention by pairwise training
Quanbin Wang and Ying Tan · 2020
Earlier work this paper cites.
Probing classifiers: Promises, shortcomings, and advances, 2021
Yonatan Belinkov · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al · 2021
Cited alongside, same era.
q 2 q^{2} : Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering
Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend · 2021
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2021
Cited alongside, same era.
Learning compact metrics for MT
Amy Pu, Hyung Won Chung, Ankur Parikh, Sebastian Gehrmann, and Thibault Sellam · 2021
Chatgpt and bard exhibit spontaneous citation fabrication during psychiatry literature search
Alessia McGowan, Yunlai Gui, Matthew Dobbs, Sophia Shuster, Matthew Cotter, Alexandria Selloni, Marianne Goodman, Agrima Srivastava, Guillermo A Cecchi, and Cheryl M Corcoran · 2023
Later among the works it cites.
LLMs confabulate not hallucinate
Beren Millidge · 2023
Later among the works it cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Chris Olah, Nelson Elhage, Neel Nanda, Catherine Schubert, Daniel Filan, et al · 2023
Later among the works it cites.
Weakly supervised detection of hallucinations in llm activations
Miriam Rateike, Celia Cintas, John Wamburu, Tanya Akumu, and Skyler Speakman · 2023
Later among the works it cites.
Vipula Rawte, Swagata Chakraborty, Agnibh Pathak, Anubhav Sarkar, SM Tonmoy, Aman Chadha, Amit P Sheth, and Amitava Das · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi · 2021
Cited alongside, same era.
QuestEval: Summarization asks for fact-based evaluation
Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari · 2021
Cited alongside, same era.
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt · 2022
Cited alongside, same era.
RED-ACE: Robust error detection for ASR using confidence embeddings
Zorik Gekhman, Dina Zverinski, Jonathan Mallinson, and Genady Beryozkin · 2022
Cited alongside, same era.
TRUE: Re-evaluating factual consistency evaluation
Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansy, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, and Yossi Matias · 2022
Cited alongside, same era.
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al · 2022
Cited alongside, same era.
SummaC: Re-visiting NLI-based models for inconsistency detection in summarization
Philippe Laban, Tobias Schnabel, Paul N. Bennett, and Marti A. Hearst · 2022
Cited alongside, same era.
Later among the works it cites.
Personality traits in large language models
Greg Serapio-García, Mustafa Safdari, Clément Crepy, Luning Sun, Stephen Fitz, Peter Romero, Marwa Abdulhai, Aleksandra Faust, and Maja Matarić · 2023
Later among the works it cites.
The curious case of hallucinatory (un)answerability: Finding truths in the hidden states of over-confident large language models
Aviv Slobodkin, Omer Goldman, Avi Caciularu, Ido Dagan, and Shauli Ravfogel · 2023
Later among the works it cites.
On early detection of hallucinations in factual question answering, 2023
Ben Snyder, Marius Moisescu, and Muhammad Bilal Zafar · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
A stitch in time saves nine: Detecting and mitigating hallucinations of llms by validating low-confidence generation, 2023
Neeraj Varshney, Wenlin Yao, Hongming Zhang, Jianshu Chen, and Dong Yu · 2023
Later among the works it cites.
Attention satisfies: A constraint-satisfaction lens on factual errors of language models
Mert Yuksekgonul, Varun Chandrasekaran, Erik Jones, Suriya Gunasekar, Ranjita Naik, Hamid Palangi, Ece Kamar, and Besmira Nushi · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency, 2023
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks · 2023
Later among the works it cites.
Truth is universal: Robust detection of lies in llms
Lennart Bürger, Fred A Hamprecht, and Boaz Nadler · 2024
Closest in time.
INSIDE: LLMs’ internal states retain the power of hallucination detection
Chao Chen, Kai Liu, Ze Chen, Yi Gu, Yue Wu, Mingyuan Tao, Zhihang Fu, and Jieping Ye · 2024
Closest in time.
Dola: Decoding by contrasting layers improves factuality in large language models
Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James R. Glass, and Pengcheng He · 2024
Closest in time.
Does fine-tuning llms on new knowledge encourage hallucinations?, 2024
Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal, Amir Feder, Roi Reichart, and Jonathan Herzig · 2024
Closest in time.
Estimating knowledge in large language models without generating a single token
Daniela Gottesman and Mor Geva · 2024
Closest in time.
Language writ large: Llms, chatgpt, grounding, meaning and understanding
Stevan Harnad · 2024
Closest in time.
Still no lie detector for language models: Probing empirical and conceptual roadblocks
Benjamin A Levinstein and Daniel A Herrmann · 2024
Closest in time.
Inference-time intervention: Eliciting truthful answers from a language model
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2024
Closest in time.
Detection-correction structure via general language model for grammatical error correction
Wei Li and Houfeng Wang · 2024
Closest in time.
Internal consistency and self-feedback in large language models: A survey
Xun Liang, Shichao Song, Zifan Zheng, Hanyu Wang, Qingchen Yu, Xunkai Li, Rong-Hua Li, Yi Wang, Zhonghao Wang, Feiyu Xiong, et al · 2024
Closest in time.
Interpreting gpt: The logit lens
nostalgebraist · 2024
Closest in time.
Constructing benchmarks and interventions for combating hallucinations in llms, 2024
Adi Simhi, Jonathan Herzig, Idan Szpektor, and Yonatan Belinkov · 2024
Closest in time.
Benchmarking hallucination in large language models based on unanswerable math word problem
Yuhong Sun, Zhangyue Yin, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Hui Zhao · 2024
Closest in time.
Pranav Narayanan Venkit, Tatiana Chakravorti, Vipul Gupta, Heidi Biggs, Mukund Srinath, Koustava Goswami, Sarah Rajtmajer, and Shomir Wilson · 2024
Closest in time.
Characterizing truthfulness in large language model generations with local intrinsic dimension
Fan Yin, Jayanth Srinivasa, and Kai-Wei Chang · 2024
Closest in time.
Can large language models faithfully express their intrinsic uncertainty in words?, 2024
Gal Yona, Roee Aharoni, and Mor Geva · 2024
Closest in time.
Inside-out: Hidden factual knowledge in llms
Zorik Gekhman, Eyal Ben David, Hadas Orgad, Eran Ofek, Yonatan Belinkov, Idan Szpector, Jonathan Herzig, and Roi Reichart · 2025
Closest in time.
Confidence improves self-consistency in llms
Amir Taubenfeld, Tom Sheffer, Eran Ofek, Amir Feder, Ariel Goldstein, Zorik Gekhman, and Gal Yona · 2025
Closest in time.