Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) excel at multi-hop reasoning in distribution, yet fail on unseen compositions, a phenomenon known as the curse of two-hop reasoning.
Guaranteed minimum-rank solutions of linear matrix equations via nuclear norm minimization
Benjamin Recht, Maryam Fazel, and Pablo A Parrilo · 2010
Earlier work this paper cites.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2019
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Earlier work this paper cites.
Locating and editing factual associations in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Earlier work this paper cites.
Measuring and narrowing the compositionality gap in language models
Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A Smith, and Mike Lewis · 2022
Earlier work this paper cites.
MuSiQue: Multihop questions via single-hop question composition
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou · 2022
Earlier work this paper cites.
Faith and fate: Limits of transformers on compositionality
Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jian, Bill Yuchen Lin, Peter West, Chandra Bhagavatula, Ronan Le Bras, and Jena D. Hwang · 2023
Earlier work this paper cites.
Understanding finetuning for factual knowledge extraction from language models
Mehran Kazemi, Sid Mittal, and Deepak Ramachandran · 2023
Earlier work this paper cites.
An investigation of llms’ inefficacy in understanding converse relations
Chengwen Qi, Bowen Li, Binyuan Hui, Bailin Wang, Jinyang Li, Jinwang Wu, and Yuanjun Laili · 2023
Earlier work this paper cites.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan · 2023
Earlier work this paper cites.
Physics of language models: part 3.1, knowledge storage and extraction
Zeyuan Allen-Zhu and Yuanzhi Li · 2024
Cited alongside, same era.
The reversal curse: LLMs trained on “a is b” fail to learn “b is a”
Lukas Berglund, Meg Tong, Maximilian Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans · 2024
Cited alongside, same era.
Hopping too late: Exploring the limitations of large language models on multi-hop queries
Eden Biran, Daniela Gottesman, Sohee Yang, Mor Geva, and Amir Globerson · 2024
Cited alongside, same era.
Break the chain: Large language models can be shortcut reasoners
Mengru Ding, Hanmeng Liu, Zhizhang Fu, Jian Song, Wenbo Xie, and Yue Zhang · 2024
Cited alongside, same era.
How to use and interpret activation patching, 2024
Stefan Heimersheim and Neel Nanda · 2024
Cited alongside, same era.
Do llms really think step-by-step in implicit reasoning?
Yijiong Yu · 2024
Later among the works it cites.
Physics of language models: Part 3.2, knowledge manipulation
Zeyuan Allen-Zhu and Yuanzhi Li · 2025
Closest in time.
Xiaoyan Bai, Itamar Pres, Yuntian Deng, Chenhao Tan, Stuart Shieber, Fernanda Viégas, Martin Wattenberg, and Andrew Lee · 2025
Closest in time.
Lessons from studying two-hop latent reasoning, 2025
Mikita Balesni, Tomek Korbak, and Owain Evans · 2025
Closest in time.
Generalization or hallucination? understanding out-of-context reasoning in transformers
Yixiao Huang, Hanlin Zhu, Tianyu Guo, Jiantao Jiao, Somayeh Sojoudi, Michael I Jordan, Stuart Russell, and Song Mei · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Interpreting key mechanisms of factual recall in transformer-based language models, 2024
Ang Lv, Yuhan Chen, Kaiyi Zhang, Yulong Wang, Lifeng Liu, Ji-Rong Wen, Jian Xie, and Rui Yan · 2024
Cited alongside, same era.
One step of gradient descent is provably the optimal in-context learner with one layer of linear self-attention
Arvind V. Mahankali, Tatsunori Hashimoto, and Tengyu Ma · 2024
Cited alongside, same era.
Team OLMo, Pete Walsh, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Shane Arora, Akshita Bhagia, Yuling Gu, Shengyi Huang, Matt Jordan, et al · 2024
Cited alongside, same era.
Grokking of implicit reasoning in transformers: A mechanistic journey to the edge of generalization
Boshi Wang, Xiang Yue, Yu Su, and Huan Sun · 2024
Cited alongside, same era.
Do large language models have compositional ability? an investigation into limitations and scalability
Zhuoyan Xu, Zhenmei Shi, and Yingyu Liang · 2024
Cited alongside, same era.
Do large language models latently perform multi-hop reasoning?
Sohee Yang, Elena Gribovskaya, Nora Kassner, Mor Geva, and Sebastian Riedel · 2024
Cited alongside, same era.
Trained transformers learn linear models in-context
Ruiqi Zhang, Spencer Frei, and Peter L. Bartlett
Cited in the paper.
Closest in time.
Implicit reasoning in transformers is reasoning through shortcuts
Tianhe Lin, Jian Xie, Siyu Yuan, and Deqing Yang · 2025
Closest in time.
An analysis for reasoning bias of language models with small initialization, 2025
Junjie Yao, Zhongwang Zhang, and Zhi-Qin John Xu · 2025
Closest in time.
How do transformers learn implicit reasoning?
Jiaran Ye, Zijun Yao, Zhidian Huang, Liangming Pan, Jinxin Liu, Yushi Bai, Amy Xin, Liu Weichuan, Xiaoyin Che, Lei Hou, and Juanzi Li · 2025
Closest in time.
Complexity control facilitates reasoning-based compositional generalization in transformers
Zhongwang Zhang, Pengxiao Lin, Zhiwei Wang, Yaoyu Zhang, and Zhi-Qin John Xu · 2025
Closest in time.
Layer-order inversion: Rethinking latent multi-hop reasoning in large language models
Xukai Liu, Ye Liu, Jipeng Zhang, Yanghai Zhang, Kai Zhang, and Qi Liu · 2026
Closest in time.