Fetching the paper…
Reading the bibliography…
This work presents a framework for assessing whether large language models (LLMs) encode more factual knowledge in their parameters than what they express in their outputs.
Aspects of the Theory of Syntax
Noam Chomsky · 1965
Earlier work this paper cites.
The “tip of the tongue” phenomenon
Roger Brown and David McNeill · 1966
Earlier work this paper cites.
The meaning and use of the area under a receiver operating characteristic (roc) curve
J. A. Hanley and B. J. McNeil · 1982
Earlier work this paper cites.
Language comprehension and production
Rebecca Treiman, Charles Clifton Jr, Antje S Meyer, and Lee H Wurm · 2003
Earlier work this paper cites.
An introduction to ROC analysis
Tom Fawcett · 2005
Earlier work this paper cites.
Wikidata: a free collaborative knowledgebase
Denny Vrandečić and Markus Krötzsch · 2014
Earlier work this paper cites.
Visualisation and ’diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure
Dieuwke Hupkes, Sara Veldhoen, and Willem H. Zuidema · 2018
Earlier work this paper cites.
Analysis methods in neural language processing: A survey
Yonatan Belinkov and James R. Glass · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
How can we know what language models know?
Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig · 2020
Earlier work this paper cites.
Are pretrained language models symbolic reasoners over knowledge?
Nora Kassner, Benno Krojer, and Hinrich Schütze · 2020
Earlier work this paper cites.
How much knowledge can you pack into the parameters of a language model?
Adam Roberts, Colin Raffel, and Noam Shazeer · 2020
Earlier work this paper cites.
Editing factual knowledge in language models
Nicola De Cao, Wilker Aziz, and Ivan Titov · 2021
Earlier work this paper cites.
Measuring and improving consistency in pretrained language models
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard H. Hovy, Hinrich Schütze, and Yoav Goldberg · 2021
Earlier work this paper cites.
BeliefBank: Adding memory to a pre-trained language model for a systematic notion of belief
Nora Kassner, Oyvind Tafjord, Hinrich Schütze, and Peter Clark · 2021
Earlier work this paper cites.
Simple entity-centric questions challenge dense retrievers
Christopher Sciavolino, Zexuan Zhong, Jinhyuk Lee, and Danqi Chen · 2021
Earlier work this paper cites.
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al · 2022
Earlier work this paper cites.
Teaching models to express their uncertainty in words
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Earlier work this paper cites.
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al · 2023
Cited alongside, same era.
The internal state of an LLM knows when it’s lying
Amos Azaria and Tom M. Mitchell · 2023
Cited alongside, same era.
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt · 2023
Cited alongside, same era.
Crawling the internal knowledge-base of language models
Roi Cohen, Mor Geva, Jonathan Berant, and Amir Globerson · 2023
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Cited alongside, same era.
Does fine-tuning LLMs on new knowledge encourage hallucinations?
Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal, Amir Feder, Roi Reichart, and Jonathan Herzig · 2024
Later among the works it cites.
Estimating knowledge in large language models without generating a single token
Daniela Gottesman and Mor Geva · 2024
Later among the works it cites.
The larger the better? improved LLM code-generation via budget reallocation
Michael Hassid, Tal Remez, Jonas Gehring, Roy Schwartz, and Yossi Adi · 2024
Later among the works it cites.
The geometry of truth: Emergent linear structure in large language model representations of true/false datasets
Samuel Marks and Max Tegmark · 2024
Later among the works it cites.
Introducing openai o1 preview, 2024
OpenAI · 2024
Later among the works it cites.
Llms know more than they show: On the intrinsic representation of llm hallucinations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Inference-time intervention: Eliciting truthful answers from a language model
Kenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister, and Martin Wattenberg · 2023
Cited alongside, same era.
Inference-time intervention: Eliciting truthful answers from a language model
Kenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister, and Martin Wattenberg · 2023
Cited alongside, same era.
Cognitive dissonance: Why do language model outputs disagree with internal representations of truthfulness?
Kevin Liu, Stephen Casper, Dylan Hadfield-Menell, and Jacob Andreas · 2023
Cited alongside, same era.
Weakly supervised detection of hallucinations in llm activations
Miriam Rateike, Celia Cintas, John Wamburu, Tanya Akumu, and Skyler Speakman · 2023
Cited alongside, same era.
Large language models encode clinical knowledge
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al · 2023
Cited alongside, same era.
Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback
Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D. Manning · 2023
Cited alongside, same era.
Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman · 2023
Cited alongside, same era.
Hadas Orgad, Michael Toker, Zorik Gekhman, Roi Reichart, Idan Szpektor, Hadas Kotek, and Yonatan Belinkov · 2024
Later among the works it cites.
Can sensitive information be deleted from llms? objectives for defending against extraction attacks
Vaidehi Patil, Peter Hase, and Mohit Bansal · 2024
Later among the works it cites.
Distinguishing ignorance from error in llm hallucinations
Adi Simhi, Jonathan Herzig, Idan Szpektor, and Yonatan Belinkov · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2024
Later among the works it cites.
On early detection of hallucinations in factual question answering
Ben Snyder, Marius Moisescu, and Muhammad Bilal Zafar · 2024
Later among the works it cites.
Unsupervised real-time hallucination detection based on the internal states of large language models
Weihang Su, Changyue Wang, Qingyao Ai, Yiran Hu, Zhijing Wu, Yujia Zhou, and Yiqun Liu · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al · 2024
Later among the works it cites.
Measuring short-form factuality in large language models
Jason Wei, Nguyen Karina, Hyung Won Chung, Yunxin Joy Jiao, Spencer Papay, Amelia Glaese, John Schulman, and William Fedus · 2024
Later among the works it cites.
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al · 2024
Later among the works it cites.
Attention satisfies: A constraint-satisfaction lens on factual errors of language models
Mert Yüksekgönül, Varun Chandrasekaran, Erik Jones, Suriya Gunasekar, Ranjita Naik, Hamid Palangi, Ece Kamar, and Besmira Nushi · 2024
Later among the works it cites.
Truthx: Alleviating hallucinations by editing large language models in truthful space
Shaolei Zhang, Tian Yu, and Yang Feng · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Confidence improves self-consistency in llms
Amir Taubenfeld, Tom Sheffer, Eran Ofek, Amir Feder, Ariel Goldstein, Zorik Gekhman, and Gal Yona · 2025
Closest in time.
Sample, scrutinize and scale: Effective inference-time search by scaling verification
Eric Zhao, Pranjal Awasthi, and Sreenivas Gollapudi · 2025
Closest in time.