Fetching the paper…
Reading the bibliography…
Recent language models generate false but plausible-sounding text with surprising frequency.
Measuring Calibration in Deep Learning
Jeremy Nixon, Mike Dusenberry, Ghassen Jerfel, Timothy Nguyen, Jeremiah Liu, Linchuan Zhang, and Dustin Tran. 2020 · 1904
Earlier work this paper cites.
Does Learning Require Memorization? A Short Tale about a Long Tail
Vitaly Feldman. 2021 · 1906
Earlier work this paper cites.
ChatGPT and Physicians’ Malpractice Risk
Michelle M. Mello and Neel Guha. 2023 · 1938
Earlier work this paper cites.
The Population Frequences of Species and the Estimation of Population Parameters
I. J. Good. 1953 · 1953
Earlier work this paper cites.
The Well-Calibrated Bayesian
A. P. Dawid. 1982 · 1982
Earlier work this paper cites.
On the Convergence Rate of Good-Turing Estimators
David McAllester and Robert E Schapire. 2000 · 2000
Earlier work this paper cites.
Concentration Inequalities for the Missing Mass and for Histogram Rule Error
David A. McAllester and Luis E. Ortiz. 2003 · 2003
Earlier work this paper cites.
Progressive Generation of Long Text with Pretrained Language Models
Bowen Tan, Zichao Yang, Maruan AI-Shedivat, Eric P. Xing, and Zhiting Hu. 2021 · 2006
Earlier work this paper cites.
Word-based predictive text entry using adaptive language models
Kumiko Tanaka-Ishii. 2007 · 2007
Earlier work this paper cites.
Controlled Hallucinations: Learning to Generate Faithfully from Noisy Data
Katja Filippova. 2020 · 2010
Earlier work this paper cites.
On the concentration of the missing mass
Daniel Berend and Aryeh Kontorovich. 2013 · 2013
Earlier work this paper cites.
Object Hallucination in Image Captioning. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (Eds.). Association for Computational Linguistics, Brussels, Belgium, 4035–4045
Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. 2018 · 2018
Earlier work this paper cites.
Calibration, Entropy Rates, and Memory in Language Models. In Proceedings of the 37th International Conference on Machine Learning . PMLR, 1089–1099
Mark Braverman, Xinyi Chen, Sham Kakade, Karthik Narasimhan, Cyril Zhang, and Yi Zhang. 2020 · 2020
Earlier work this paper cites.
On Faithfulness and Factuality in Abstractive Summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computational Linguistics, Online, 1906–1919
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020 · 2020
Cited alongside, same era.
Truthful AI: Developing and governing AI that does not lie
Owain Evans, Owen Cotton-Barratt, Lukas Finnveden, Adam Bales, Avital Balwit, Peter Wills, Luca Righetti, and William Saunders. 2021 · 2021
Cited alongside, same era.
How Can We Know When Language Models Know? On the Calibration of Language Models for Question Answering
Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021 · 2021
Cited alongside, same era.
Calibrate Before Use: Improving Few-shot Performance of Language Models. In Proceedings of the 38th International Conference on Machine Learning . PMLR, 12697–12706
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021 · 2021
Cited alongside, same era.
Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence
Joseph R. Jr. Biden. 2023 · 2023
Closest in time.
Loss minimization yields multicalibration for large neural networks
Jarosław Błasiok, Parikshit Gopalan, Lunjia Hu, Adam Tauman Kalai, and Preetum Nakkiran. 2023 · 2023
Closest in time.
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2023 · 2023
Closest in time.
Survey of Hallucination in Natural Language Generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023 · 2023
Closest in time.
SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improving language models by retrieving from trillions of tokens
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Chris Jones, Albin Cassirer, Andy Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, and Laurent Sifre. 2022 · 2022
Cited alongside, same era.
PaLM: Scaling Language Modeling with Pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022 · 2022
Cited alongside, same era.
On the Origin of Hallucinations in Conversational Models: Is it the Datasets or the Models? arXiv
Nouha Dziri, Sivan Milton, Mo Yu, Osmar Zaiane, and Siva Reddy. 2022 · 2022
Cited alongside, same era.
Language Models (Mostly) Know What They Know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom Brown, Jack Clark, Nicholas Joseph, Ben Mann, Sam McCandlish, Chris Olah, and Jared Kaplan. 2022 · 2022
Cited alongside, same era.
TruthfulQA: Measuring How Models Mimic Human Falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, Dublin, Ireland, 3214–3252
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Cited alongside, same era.
Do Language Models Know When They’re Hallucinating References?
Ayush Agrawal, Mirac Suzgun, Lester Mackey, and Adam Tauman Kalai. 2023 · 2023
Cited alongside, same era.
Characterizing Attribution and Fluency Tradeoffs for Retrieval-Augmented Large Language Models
Renat Aksitov, Chung-Ching Chang, David Reitter, Siamak Shakeri, and Yunhsuan Sung. 2023 · 2023
Cited alongside, same era.
The Internal State of an LLM Knows When its Lying
Amos Azaria and Tom Mitchell. 2023 · 2023
Cited alongside, same era.
Potsawee Manakul, Adian Liusie, and Mark J. F. Gales. 2023 · 2023
Closest in time.
Definition of HALLUCINATION
Merriam-Webster. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
SlimPajama-DC: Understanding Data Combinations for LLM Training
Zhiqiang Shen, Tianhua Tao, Liqun Ma, Willie Neiswanger, Zhengzhong Liu, Hongyi Wang, Bowen Tan, Joel Hestness, Natalia Vassilieva, Daria Soboleva, and Eric Xing. 2023 · 2023
Closest in time.
Humiliated lawyers fined $5,000 for submitting ChatGPT hallucinations in court: ‘I heard about this new site, which I falsely assumed was, like, a super search engine’
Rachel Shin. 2023 · 2023
Closest in time.
FreshLLMs: Refreshing Large Language Models with Search Engine Augmentation
Tu Vu, Mohit Iyyer, Xuezhi Wang, Noah Constant, Jerry Wei, Jason Wei, Chris Tar, Yun-Hsuan Sung, Denny Zhou, Quoc Le, and Thang Luong. 2023 · 2023
Closest in time.
When A.I. Chatbots Hallucinate
Karen Weise and Cade Metz. 2023 · 2023
Closest in time.
Understanding and Detecting Hallucinations in Neural Machine Translation via Model Introspection
Weijia Xu, Sweta Agrawal, Eleftheria Briakou, Marianna J. Martindale, and Marine Carpuat. 2023 · 2023
Closest in time.