Fetching the paper…
Reading the bibliography…
What are the root causes of hallucinations in large language models (LLMs)? We use Communication Complexity to prove that the Transformer layer is incapable of composing functions (e.g., identify a grandparent of a person in a genealogy) if the domains of the functions are large enough; we show through examples that this inability is already empirically present when the domains are quite small.
Communication complexity
Christos H Papadimitriou and Michael Sipser · 1982
Earlier work this paper cites.
Rounds in communication complexity revisited
Noam Nisan and Avi Widgerson · 1991
Earlier work this paper cites.
Computational complexity
Christos H Papadimitriou · 1993
Earlier work this paper cites.
Communication complexity
Eyal Kushilevitz and Noam Nisan · 1996
Earlier work this paper cites.
On quantum and probabilistic communication: Las vegas and one-way protocols
Hartmut Klauck · 2000
Earlier work this paper cites.
Computational complexity: a modern approach
Sanjeev Arora and Boaz Barak · 2009
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
How can self-attention networks recognize dyck-n languages?
Javid Ebrahimi, Dhruv Gelda, and Wei Zhang · 2020
Earlier work this paper cites.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn · 2020
Cited alongside, same era.
Pointer chasing via triangular discrimination
Amir Yehudayoff · 2020
Cited alongside, same era.
Self-attention networks can process bounded hierarchical languages
Shunyu Yao, Binghui Peng, Christos Papadimitriou, and Karthik Narasimhan · 2021
Cited alongside, same era.
Memory bounds for continual learning
Xi Chen, Christos Papadimitriou, and Binghui Peng · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
Faith and fate: Limits of transformers on compositionality
Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jiang, Bill Yuchen Lin, Peter West, Chandra Bhagavatula, Ronan Le Bras, Jena D Hwang, et al · 2023
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung · 2023
Later among the works it cites.
Calibrated language models must hallucinate
Adam Tauman Kalai and Santosh S Vempala · 2023
Later among the works it cites.
The expresssive power of transformers with chain of thought
William Merrill and Ashish Sabharwal · 2023
Later among the works it cites.
The parallelism tradeoff: Limitations of log-precision transformers
William Merrill and Ashish Sabharwal · 2023
Later among the works it cites.
R Thomas McCoy, Shunyu Yao, Dan Friedman, Matthew Hardy, and Thomas L Griffiths · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Towards revealing the mystery behind chain of thought: a theoretical perspective
Guhao Feng, Yuntian Gu, Bohang Zhang, Haotian Ye, Di He, and Liwei Wang · 2023
Cited alongside, same era.
Holy Bible, Matthew 1, 1–17
Cited in the paper.
Later among the works it cites.
Representational strengths and limitations of transformers
Clayton Sanford, Daniel Hsu, and Matus Telgarsky · 2023
Later among the works it cites.
Mitigating large language model hallucinations via autonomous knowledge graph-based retrofitting
Xinyan Guan, Yanjiang Liu, Hongyu Lin, Yaojie Lu, Ben He, Xianpei Han, and Le Sun · 2024
Closest in time.