Fetching the paper…
Reading the bibliography…
Hallucination is often regarded as a major impediment for using large language models (LLMs), especially for knowledge-intensive tasks.
Word association norms, mutual information, and lexicography
Kenneth Church and Patrick Hanks · 1990
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer · 2015
Earlier work this paper cites.
Information-theoretic analysis of generalization capability of learning algorithms
Aolin Xu and Maxim Raginsky · 2017
Earlier work this paper cites.
Chaining mutual information and tightening generalization bounds
Amir Asadi, Emmanuel Abbe, and Sergio Verdú · 2018
Earlier work this paper cites.
Foundations of machine learning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar · 2018
Earlier work this paper cites.
Generalization error bounds for noisy, iterative algorithms
Ankit Pensia, Varun Jog, and Po-Ling Loh · 2018
Earlier work this paper cites.
Learning imbalanced datasets with label-distribution-aware margin loss
Kaidi Cao, Colin Wei, Adrien Gaidon, Nikos Arechiga, and Tengyu Ma · 2019
Earlier work this paper cites.
Information-theoretic generalization bounds for sgld via data-dependent estimates
Jeffrey Negrea, Mahdi Haghifam, Gintare Karolina Dziugaite, Ashish Khisti, and Daniel M Roy · 2019
Earlier work this paper cites.
How much does your data exploration overfit? controlling bias via information usage
Daniel Russo and James Zou · 2019
Earlier work this paper cites.
Revisiting challenges in data-to-text generation with fact grounding
Hongmin Wang · 2019
Earlier work this paper cites.
Data-dependent sample complexity of deep neural networks via lipschitz augmentation
Colin Wei and Tengyu Ma · 2019
Earlier work this paper cites.
Understanding why neural networks generalize well through gsnr of parameters
Jinlong Liu, Guoqing Jiang, Yunzhi Bai, Ting Chen, and Huayan Wang · 2020
Earlier work this paper cites.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald · 2020
Earlier work this paper cites.
Long-tail learning via logit adjustment
Aditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain, Andreas Veit, and Sanjiv Kumar · 2020
Earlier work this paper cites.
ToTTo: A controlled table-to-text generation dataset
Ankur Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, and Dipanjan Das · 2020
Earlier work this paper cites.
Reasoning about generalization via conditional mutual information
Thomas Steinke and Lydia Zakynthinou · 2020
Earlier work this paper cites.
Information-theoretic generalization bounds for black-box learning algorithms
Hrayr Harutyunyan, Maxim Raginsky, Greg Ver Steeg, and Aram Galstyan · 2021
Earlier work this paper cites.
Label-imbalanced and group-sensitive classification under overparameterization
Ganesh Ramachandra Kini, Orestis Paraskevas, Samet Oymak, and Christos Thrampoulidis · 2021
Earlier work this paper cites.
The curious case of hallucinations in neural machine translation
Vikas Raunak, Arul Menezes, and Marcin Junczys-Dowmunt · 2021
Earlier work this paper cites.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki · 2021
Earlier work this paper cites.
#HowYouTagTweets: Learning user hashtagging preferences via personalized topic attention
Yuji Zhang, Yubo Zhang, Chunpu Xu, Jing Li, Ziyan Jiang, and Baolin Peng · 2021
Earlier work this paper cites.
On the origin of hallucinations in conversational models: Is it the datasets or the models?
Nouha Dziri, Sivan Milton, Mo Yu, Osmar Zaiane, and Siva Reddy · 2022
Earlier work this paper cites.
Language models (mostly) know what they know, 2022
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom Brown, Jack Clark, Nicholas Joseph, Ben Mann, Sam McCandlish, Chris Olah, and Jared Kaplan · 2022
Earlier work this paper cites.
Large language models with controllable working memory, 2022
Daliang Li, Ankit Singh Rawat, Manzil Zaheer, Xin Wang, Michal Lukasik, Andreas Veit, Felix Yu, and Sanjiv Kumar · 2022
Earlier work this paper cites.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Earlier work this paper cites.
Streamingqa: A benchmark for adaptation to new knowledge over time in question answering models
Adam Livska, Tom’avs Kovcisk’y, Elena Gribovskaya, Tayfun Terzi, Eren Sezener, Devang Agrawal, Cyprien de Masson d’Autume, Tim Scholtes, Manzil Zaheer, Susannah Young, Ellen Gilsenan-McMahon, Sophia Austin, Phil Blunsom, and Angeliki Lazaridou · 2022
Earlier work this paper cites.
Time waits for no one! analysis and challenges of temporal misalignment
Kelvin Luu, Daniel Khashabi, Suchin Gururangan, Karishma Mandyam, and Noah A. Smith · 2022
Earlier work this paper cites.
Locating and editing factual associations in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Earlier work this paper cites.
Discovering language model behaviors with model-written evaluations, 2022
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Ben Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan · 2022
Earlier work this paper cites.
Read before generate! faithful long form question answering with machine reading
Dan Su, Xiaoguang Li, Jindi Zhang, Lifeng Shang, Xin Jiang, Qun Liu, and Pascale Fung · 2022
Cited alongside, same era.
Yuji Zhang and Jing Li · 2022
Cited alongside, same era.
Self-rag: Learning to retrieve, generate, and critique through self-reflection, 2023
Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi · 2023
Cited alongside, same era.
The internal state of an llm knows when it’s lying, 2023
Amos Azaria and Tom Mitchell · 2023
Cited alongside, same era.
Purr: Efficiently editing language model hallucinations by denoising language model corruptions, 2023
Anthony Chen, Panupong Pasupat, Sameer Singh, Hongrae Lee, and Kelvin Guu · 2023
Cited alongside, same era.
On early detection of hallucinations in factual question answering, 2023
Ben Snyder, Marius Moisescu, and Muhammad Bilal Zafar · 2023
Later among the works it cites.
Towards fair financial services for all: A temporal GNN approach for individual fairness on transaction networks
Zixing Song, Yuji Zhang, and Irwin King · 2023
Later among the works it cites.
Towards fair financial services for all: A temporal gnn approach for individual fairness on transaction networks
Zixing Song, Yuji Zhang, and Irwin King · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Don’t believe everything you read: Enhancing summarization interpretability through automatic identification of hallucinations in large language models, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dola: Decoding by contrasting layers improves factuality in large language models, 2023
Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James Glass, and Pengcheng He · 2023
Cited alongside, same era.
Generalization bounds using data-dependent fractal dimensions
Benjamin Dupuis, George Deligiannidis, and Umut Simsekli · 2023
Cited alongside, same era.
Halo: Estimation and reduction of hallucinations in open-source weak large language models
Mohamed Elaraby, Mengyin Lu, Jacob Dunn, Xueying Zhang, Yu Wang, and Shizhu Liu · 2023
Cited alongside, same era.
Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, et al · 2023
Cited alongside, same era.
Large language models cannot self-correct reasoning yet
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou · 2023
Cited alongside, same era.
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al · 2023
Cited alongside, same era.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung · 2023
Cited alongside, same era.
Priyesh Vakharia, Devavrat Joshi, Meenal Chavan, Dhananjay Sonawane, Bhrigu Garg, Parsa Mazaheri, and Ian Lane · 2023
Later among the works it cites.
A stitch in time saves nine: Detecting and mitigating hallucinations of llms by validating low-confidence generation, 2023
Neeraj Varshney, Wenlin Yao, Hongming Zhang, Jianshu Chen, and Dong Yu · 2023
Later among the works it cites.
Reducing llm hallucinations using epistemic neural networks, 2023
Shreyas Verma, Kien Tran, Yusuf Ali, and Guangyu Min · 2023
Later among the works it cites.
Simple synthetic data reduces sycophancy in large language models, 2023
Jerry Wei, Da Huang, Yifeng Lu, Denny Zhou, and Quoc V. Le · 2023
Later among the works it cites.
Adaptive chameleon or stubborn sloth: Unraveling the behavior of large language models in knowledge clashes, 2023
Jian Xie, Kai Zhang, Jiangjie Chen, Renze Lou, and Yu Su · 2023
Later among the works it cites.
Adept: A debiasing prompt framework
Ke Yang, Charles Yu, Yi R. Fung, Manling Li, and Heng Ji · 2023
Later among the works it cites.
Cognitive mirage: A review of hallucinations in large language models, 2023
Hongbin Ye, Tong Liu, Aijia Zhang, Wei Hua, and Weiqiang Jia · 2023
Later among the works it cites.
Do large language models know what they don’t know?, 2023
Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Xuanjing Huang · 2023
Later among the works it cites.
Automatic evaluation of attribution by large language models, 2023
Xiang Yue, Boshi Wang, Kai Zhang, Ziru Chen, Yu Su, and Huan Sun · 2023
Later among the works it cites.
How language model hallucinations can snowball, 2023
Muru Zhang, Ofir Press, William Merrill, Alisa Liu, and Noah A. Smith · 2023
Later among the works it cites.
Mitigating language model hallucination with interactive question-knowledge alignment, 2023
Shuo Zhang, Liangming Pan, Junzhou Zhao, and William Yang Wang · 2023
Later among the works it cites.
Alleviating hallucinations of large language models through induced hallucinations, 2023
Yue Zhang, Leyang Cui, Wei Bi, and Shuming Shi · 2023
Later among the works it cites.
Siren’s song in the ai ocean: A survey on hallucination in large language models, 2023
Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, Longyue Wang, Anh Tuan Luu, Wei Bi, Freda Shi, and Shuming Shi · 2023
Later among the works it cites.
Vibe: Topic-driven temporal adaptation for twitter classification
Yuji Zhang, Jing Li, and Wenjie Li · 2023
Later among the works it cites.
Beyond hallucinations: Enhancing lvlms through hallucination-aware direct preference optimization, 2023
Zhiyuan Zhao, Bin Wang, Linke Ouyang, Xiaoyi Dong, Jiaqi Wang, and Conghui He · 2023
Later among the works it cites.
Evedit: Event-based knowledge editing with deductive editing boundaries
Jiateng Liu, Pengfei Yu, Yuji Zhang, Sha Li, Zixuan Zhang, and Heng Ji · 2024
Closest in time.
Prejudice and volatility: A statistical framework for measuring social discrimination in large language models, 2024
Y Liu, K Yang, Z Qi, X Liu, Y Yu, and C Zhai · 2024
Closest in time.
A survey on vision-language-action models for embodied ai
Yueen Ma, Zixing Song, Yuzheng Zhuang, Jianye Hao, and Irwin King · 2024
Closest in time.
Inverse scaling: When bigger isn’t better, 2024
Ian R. McKenzie, Alexander Lyzhov, Michael Pieler, Alicia Parrish, Aaron Mueller, Ameya Prabhu, Euan McLean, Aaron Kirtland, Alexis Ross, Alisa Liu, Andrew Gritsevskiy, Daniel Wurgaft, Derik Kauffman, Gabriel Recchia, Jiacheng Liu, Joe Cavanagh, Max Weiss, Sicong Huang, The Floating Droid, Tom Tseng, Tomasz Korbak, Xudong Shen, Yuhui Zhang, Zhengping Zhou, Najoung Kim, Samuel R. Bowman, and Ethan Perez · 2024
Closest in time.
Dolma: An open corpus of three trillion tokens for language model pretraining research
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, et al · 2024
Closest in time.
A comprehensive survey of hallucination mitigation techniques in large language models, 2024
S. M Towhidul Islam Tonmoy, S M Mehedi Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das · 2024
Closest in time.
Sequence-level certainty reduces hallucination in knowledge-grounded dialogue generation, 2024
Yixin Wan, Fanyou Wu, Weijie Xu, and Srinivasan H. Sengamedu · 2024
Closest in time.
A unified generalization analysis of re-weighting and logit-adjustment for imbalanced learning
Zitai Wang, Qianqian Xu, Zhiyong Yang, Yuan He, Xiaochun Cao, and Qingming Huang · 2024
Closest in time.
R-tuning: Instructing large language models to say ‘i don’t know’, 2024
Hanning Zhang, Shizhe Diao, Yong Lin, Yi R. Fung, Qing Lian, Xingyao Wang, Yangyi Chen, Heng Ji, and Tong Zhang · 2024
Closest in time.