Fetching the paper…
Reading the bibliography…
Recent advances in large language models (LLMs) have popularized the chain-of-thought (CoT) paradigm, in which models produce explicit reasoning steps in natural language.
Working memory: looking back and looking forward
Alan Baddeley · 2003
Earlier work this paper cites.
Functional specificity for high-level linguistic processing in the human brain
Evelina Fedorenko, Michael K. Behr, and Nancy Kanwisher · 2011
Earlier work this paper cites.
Thought beyond language: Neural dissociation of algebra and natural language
Martin M. Monti, Lawrence M. Parsons, and Daniel N. Osherson · 2012
Earlier work this paper cites.
A distinct cortical network for mathematical knowledge in the human brain
Marie Amalric and Stanislas Dehaene · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anna Bakhtin, Yuxiang Wu, and Alexander Miller · 2019
Cited alongside, same era.
A mathematical framework for transformer circuits, 2021
Elhage et al · 2021
Cited alongside, same era.
Scaling laws under the microscope: Predicting transformer performance from small scale experiments
Maor Ivgi, Yair Carmon, and Jonathan Berant · 2022
Cited alongside, same era.
nanogpt: The simplest, fastest repository for training/finetuning medium-sized gpts, 2023
Andrej Karpathy · 2023
Cited alongside, same era.
What language model architecture and pretraining objective work best for zero-shot generalization?
Jason Wei et al
Cited in the paper.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei et al
Cited in the paper.
Language is primarily a tool for communication rather than thought
Evelina Fedorenko, Steven T. Piantadosi, and Edward A. F. Gibson · 2024
Later among the works it cites.
Training large language models to reason in a continuous latent space, 2024
Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, and Yuandong Tian · 2024
Later among the works it cites.
Scaling up test-time compute with latent reasoning: A recurrent depth approach, 2025
Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian R. Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, and Tom Goldstein · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…