Fetching the paper…
Reading the bibliography…
Neural language models (LMs) can be used to evaluate the truth of factual statements in two ways: they can be either queried for statement probabilities, or probed for internal representations of truthfulness.
The definition of lying and deception
James Edwin Mahon. 2008 · 2008
Earlier work this paper cites.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Identifying fluently inadequate output in neural and statistical machine translation
Marianna Martindale, Marine Carpuat, Kevin Duh, and Paul McNamee. 2019 · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
Probing the probing paradigm: Does probing accuracy entail task relevance?
Abhilasha Ravichander, Yonatan Belinkov, and Eduard Hovy. 2020 · 2020
Earlier work this paper cites.
Information-theoretic probing with minimum description length
Elena Voita and Ivan Titov. 2020 · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al. 2021 · 2021
Earlier work this paper cites.
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg. 2021 · 2021
Earlier work this paper cites.
Truthful AI: Developing and governing AI that does not lie
Owain Evans, Owen Cotton-Barratt, Lukas Finnveden, Adam Bales, Avital Balwit, Peter Wills, Luca Righetti, and William Saunders. 2021 · 2021
Cited alongside, same era.
The low-dimensional linear geometry of contextualized word representations
Evan Hernandez and Jacob Andreas. 2021 · 2021
Cited alongside, same era.
CREAK: A dataset for commonsense reasoning over entity knowledge
Yasumasa Onoe, Michael JQ Zhang, Eunsol Choi, and Greg Durrett. 2021 · 2021
Cited alongside, same era.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki. 2021 · 2021
Cited alongside, same era.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov. 2022 · 2022
Cited alongside, same era.
Discovering latent knowledge in language models without supervision
Linear adversarial concept erasure
Shauli Ravfogel, Michael Twiton, Yoav Goldberg, and Ryan D Cotterell. 2022 · 2022
Later among the works it cites.
Neural theory-of-mind? On the limits of social intelligence in large lms
Maarten Sap, Ronan Le Bras, Daniel Fried, and Yejin Choi. 2022 · 2022
Later among the works it cites.
Talking about large language models
Murray Shanahan. 2022 · 2022
Later among the works it cites.
Mirages: On anthropomorphism in dialogue systems
Gavin Abercrombie, Amanda Cercas Curry, Tanvi Dinkar, and Zeerak Talat. 2023 · 2023
Closest in time.
The internal state of an LLM knows when it’s lying
Amos Azaria and Tom Mitchell. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. 2022 · 2022
Cited alongside, same era.
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield Dodds, Nova DasSarma, Eli Tran-Johnson, et al. 2022 · 2022
Cited alongside, same era.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Cited alongside, same era.
Reducing conversational agents’ overconfidence through linguistic calibration
Sabrina J Mielke, Arthur Szlam, Emily Dinan, and Y-Lan Boureau. 2022 · 2022
Cited alongside, same era.
Why ChatGPT and Bing Chat are so good at making things up
Benj Edwards. 2023 · 2023
Closest in time.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023 · 2023
Closest in time.
Samuel Marks and Max Tegmark. 2023 · 2023
Closest in time.
Representation engineering: A top-down approach to AI transparency
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al. 2023 · 2023
Closest in time.