Fetching the paper…
Reading the bibliography…
The hallucination problem of Large Language Models (LLMs) significantly limits their reliability and trustworthiness.
Memory and the feeling-of-knowing experience
Julian T Hart. 1965 · 1965
Earlier work this paper cites.
Towards few-shot fact-checking via perplexity
Nayeon Lee, Yejin Bang, Andrea Madotto, and Pascale Fung. 2021 · 1981
Earlier work this paper cites.
Metamemory: A theoretical framework and new findings
Thomas O Nelson. 1990 · 1990
Earlier work this paper cites.
Monitoring one’s own knowledge during study: A cue-utilization approach to judgments of learning
Asher Koriat. 1997 · 1997
Earlier work this paper cites.
Estimating mutual information
Alexander Kraskov, Harald Stögbauer, and Peter Grassberger. 2004 · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Different varieties of uncertainty in human decision-making
Amy R Bland and Alexandre Schaefer. 2012 · 2012
Earlier work this paper cites.
The neural basis of metacognitive ability
Stephen M Fleming and Raymond J Dolan. 2012 · 2012
Earlier work this paper cites.
Metacognition in human decision-making: confidence and error monitoring
Nick Yeung and Christopher Summerfield. 2012 · 2012
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. 2017 · 2017
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for squad
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
Out-of-domain detection based on generative adversarial network
Seonghan Ryu, Sangjun Koo, Hwanjo Yu, and Gary Geunbae Lee. 2018 · 2018
Earlier work this paper cites.
Analysis methods in neural language processing: A survey
Yonatan Belinkov and James Glass. 2019 · 2019
Earlier work this paper cites.
Out-of-domain detection for low-resource text classification tasks
Ming Tan, Yang Yu, Haoyu Wang, Dakuo Wang, Saloni Potdar, Shiyu Chang, and Mo Yu. 2019 · 2019
Earlier work this paper cites.
Out-of-domain detection for natural language understanding in dialog systems
Yinhe Zheng, Guanyi Chen, and Minlie Huang. 2020 · 2020
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. 2021 · 2021
Earlier work this paper cites.
Internal language model estimation for domain-adaptive end-to-end speech recognition
Zhong Meng, Sarangarajan Parthasarathy, Eric Sun, Yashesh Gaur, Naoyuki Kanda, Liang Lu, Xie Chen, Rui Zhao, Jinyu Li, and Yifan Gong. 2021 · 2021
Earlier work this paper cites.
Questeval: Summarization asks for fact-based evaluation
Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari. 2021 · 2021
Earlier work this paper cites.
On hallucination and predictive uncertainty in conditional language generation
Yijun Xiao and William Yang Wang. 2021 · 2021
Cited alongside, same era.
Generalized out-of-distribution detection: A survey
Jingkang Yang, Kaiyang Zhou, Yixuan Li, and Ziwei Liu. 2021 · 2021
Cited alongside, same era.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov. 2022 · 2022
Cited alongside, same era.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Andrea Madotto, and Pascale Fung. 2022 · 2022
Cited alongside, same era.
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al. 2022 · 2022
Cited alongside, same era.
Factscore: Fine-grained atomic evaluation of factual precision in long form text generation
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023 · 2023
Later among the works it cites.
The curious case of hallucinatory (un) answerability: Finding truths in the hidden states of over-confident large language models
Aviv Slobodkin, Omer Goldman, Avi Caciularu, Ido Dagan, and Shauli Ravfogel. 2023 · 2023
Later among the works it cites.
On early detection of hallucinations in factual question answering
Ben Snyder, Marius Moisescu, and Muhammad Bilal Zafar. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Entity cloze by date: What LMs know about unseen entities
Yasumasa Onoe, Michael Zhang, Eunsol Choi, and Greg Durrett. 2022 · 2022
Cited alongside, same era.
Uncertainty quantification with pre-trained language models: A large-scale empirical analysis
Yuxin Xiao, Paul Pu Liang, Umang Bhatt, Willie Neiswanger, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2022 · 2022
Cited alongside, same era.
Hidden state variability of pretrained language models can guide computation reduction for transfer learning
Shuo Xie, Jiahao Qiu, Ankita Pasad, Li Du, Qing Qu, and Hongyuan Mei. 2022 · 2022
Cited alongside, same era.
The internal state of an llm knows when it’s lying
Amos Azaria and Tom Mitchell. 2023 · 2023
Cited alongside, same era.
A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity
Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, and Pascale Fung. 2023 · 2023
Cited alongside, same era.
Language models can explain neurons in language models
Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders. 2023 · 2023
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah. 2023 · 2023
Cited alongside, same era.
Yile Wang, Peng Li, Maosong Sun, and Yang Liu. 2023 · 2023
Later among the works it cites.
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms
Miao Xiong, Zhiyuan Hu, Xinyang Lu, YIFEI LI, Jie Fu, Junxian He, and Bryan Hooi. 2023 · 2023
Later among the works it cites.
Alcuna: Large language models meet new knowledge
Xunjian Yin, Baizhou Huang, and Xiaojun Wan. 2023a · 2023
Later among the works it cites.
Do large language models know what they don’t know?
Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Xuanjing Huang. 2023b · 2023
Later among the works it cites.
Two birds one stone: Dynamic ensemble for ood intent classification
Yunhua Zhou, Jianqiang Yang, Pengyu Wang, and Xipeng Qiu. 2023 · 2023
Later among the works it cites.
Distinguishing the knowable from the unknowable with language models
Gustaf Ahdritz, Tian Qin, Nikhil Vyas, Boaz Barak, and Benjamin L Edelman. 2024 · 2024
Closest in time.
INSIDE: LLMs’ internal states retain the power of hallucination detection
Chao Chen, Kai Liu, Ze Chen, Yi Gu, Yue Wu, Mingyuan Tao, Zhihang Fu, and Jieping Ye. 2024 · 2024
Closest in time.
Estimating knowledge in large language models without generating a single token
Daniela Gottesman and Mor Geva. 2024 · 2024
Closest in time.
Anah: Analytical annotation of hallucinations in large language models
Ziwei Ji, Yuzhe Gu, Wenwei Zhang, Chengqi Lyu, Dahua Lin, and Kai Chen. 2024 · 2024
Closest in time.
Latesteval: Addressing data contamination in language model evaluation through dynamic and time-sensitive test construction
Yucheng Li, Frank Guerin, and Chenghua Lin. 2024 · 2024
Closest in time.
Uncertainty estimation and quantification for llms: A simple supervised approach
Linyu Liu, Yu Pan, Xiaocheng Li, and Guanting Chen. 2024 · 2024
Closest in time.
Unsupervised real-time hallucination detection based on the internal states of large language models
Weihang Su, Changyue Wang, Qingyao Ai, Yiran Hu, Zhijing Wu, Yujia Zhou, and Yiqun Liu. 2024 · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan. 2024 · 2024
Closest in time.
Extracting concepts from gpt-4
Jeffrey Wu, Leo Gao, Tom Dupre la Tour, and Henk Tillman. 2024 · 2024
Closest in time.