Fetching the paper…
Reading the bibliography…
Modern language models generate chain-of-thought traces by autoregressively sampling tokens from a finite vocabulary.
Latent dirichlet allocation
D. M. Blei, A. Y. Ng, and M. I. Jordan · 2003
Earlier work this paper cites.
Learning to summarize from human feedback, 2022
N. Stiennon, L. Ouyang, J. Wu, D. M. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. Christiano · 2009
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks, 2015
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer · 2015
Earlier work this paper cites.
On learning distributions from their samples
S. Kamath, A. Orlitsky, D. Pichapati, and A. T. Suresh · 2015
Earlier work this paper cites.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, and et al · 2017
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Earlier work this paper cites.
Glu variants improve transformer
N. Shazeer · 2020
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
M. Nye, A. J. Andreassen, G. Gur-Ari, H. Michalewski, J. Austin, D. Bieber, D. Dohan, A. Lewkowycz, M. Bosma, D. Luan, et al · 2021
Earlier work this paper cites.
Large language models are zero-shot reasoners
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Earlier work this paper cites.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought
A. Saparov and H. He · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
Implicit chain of thought reasoning via knowledge distillation
Y. Deng, K. Prasad, R. Fernandez, P. Smolensky, V. Chaudhary, and S. Shieber · 2023
Earlier work this paper cites.
Faith and fate: Limits of transformers on compositionality, 2023
N. Dziri, X. Lu, M. Sclar, X. L. Li, L. Jiang, B. Y. Lin, P. West, C. Bhagavatula, R. L. Bras, J. D. Hwang, S. Sanyal, S. Welleck, X. Ren, A. Ettinger, Z. Harchaoui, and Y. Choi · 2023
Earlier work this paper cites.
Towards revealing the mystery behind chain of thought: A theoretical perspective
G. Feng, B. Zhang, Y. Gu, H. Ye, D. He, and L. Wang · 2023
Earlier work this paper cites.
Looped transformers as programmable computers
A. Giannou, S. Rajput, J.-Y. Sohn, K. Lee, J. D. Lee, and D. Papailiopoulos · 2023
Cited alongside, same era.
Progress measures for grokking via mechanistic interpretability, 2023
N. Nanda, L. Chan, T. Lieberum, J. Smith, and J. Steinhardt · 2023
Cited alongside, same era.
Challenging BIG-bench tasks and whether chain-of-thought can solve them
M. Suzgun, N. Scales, N. Schärli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. Le, E. Chi, D. Zhou, and J. Wei · 2023
Cited alongside, same era.
Tree of thoughts: Deliberate problem solving with large language models
S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao, and K. Narasimhan · 2023
Cited alongside, same era.
The pitfalls of next-token prediction
G. Bachmann and V. Nagarajan · 2024
Cited alongside, same era.
Compressed chain of thought: Efficient reasoning through dense representations
Distributional reasoning in llms: Parallel reasoning processes in multi-hop reasoning
Y. Shalev, A. Feder, and A. Goldstein · 2024
Later among the works it cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models, 2024
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. K. Li, Y. Wu, and D. Guo · 2024
Later among the works it cites.
Limits of transformer language models on learning to compose algorithms, 2024
J. Thomm, G. Camposampiero, A. Terzic, M. Hersche, B. Schölkopf, and A. Rahimi · 2024
Later among the works it cites.
Do large language models latently perform multi-hop reasoning?
S. Yang, E. Gribovskaya, N. Kassner, M. Geva, and S. Riedel · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Cheng and B. Van Durme · 2024
Cited alongside, same era.
From explicit cot to implicit cot: Learning to internalize cot step by step
Y. Deng, Y. Choi, and S. Shieber · 2024
Cited alongside, same era.
Better & faster large language models via multi-token prediction
F. Gloeckle, B. Y. Idrissi, B. Rozière, D. Lopez-Paz, and G. Synnaeve · 2024
Cited alongside, same era.
Think before you speak: Training language models with pause tokens
S. Goyal, Z. Ji, A. S. Rawat, A. K. Menon, S. Kumar, and V. Nagarajan · 2024
Cited alongside, same era.
Training large language models to reason in a continuous latent space
S. Hao, S. Sukhbaatar, D. Su, X. Li, Z. Hu, J. Weston, and Y. Tian · 2024
Cited alongside, same era.
From self-attention to markov models: Unveiling the dynamics of generative transformers
M. E. Ildiz, Y. Huang, Y. Li, A. S. Rawat, and S. Oymak · 2024
Cited alongside, same era.
A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, et al · 2024
Cited alongside, same era.
P. Yu, J. Xu, J. Weston, and I. Kulikov · 2024
Later among the works it cites.
L1: Controlling how long a reasoning model thinks with reinforcement learning
P. Aggarwal and S. Welleck · 2025
Closest in time.
Scaling up test-time compute with latent reasoning: A recurrent depth approach
J. Geiping, S. McLeish, N. Jain, J. Kirchenbauer, S. Singh, B. R. Bartoldson, B. Kailkhura, A. Bhatele, and T. Goldstein · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
Reasoning with latent thoughts: On the power of looped transformers
N. Saunshi, N. Dikkala, Z. Li, S. Kumar, and S. J. Reddi · 2025
Closest in time.
Codi: Compressing chain-of-thought into continuous space via self-distillation
Z. Shen, H. Yan, L. Zhang, Z. Hu, Y. Du, and Y. He · 2025
Closest in time.
Stop overthinking: A survey on efficient reasoning for large language models
Y. Sui, Y.-N. Chuang, G. Wang, J. Zhang, T. Zhang, J. Yuan, H. Liu, A. Wen, S. Zhong, H. Chen, et al · 2025
Closest in time.
Everything everywhere all at once: Llms can in-context learn multiple tasks in superposition
Z. Xiong, Z. Cai, J. Cooper, A. Ge, V. Papageorgiou, Z. Sifakis, A. Giannou, Z. Lin, L. Yang, S. Agarwal, et al · 2025
Closest in time.
Dapo: An open-source llm reinforcement learning system at scale, 2025
Q. Yu, Z. Zhang, R. Zhu, Y. Yuan, X. Zuo, Y. Yue, T. Fan, G. Liu, L. Liu, X. Liu, H. Lin, Z. Lin, B. Ma, G. Sheng, Y. Tong, C. Zhang, M. Zhang, W. Zhang, H. Zhu, J. Zhu, J. Chen, J. Chen, C. Wang, H. Yu, W. Dai, Y. Song, X. Wei, H. Zhou, J. Liu, W.-Y. Ma, Y.-Q. Zhang, L. Yan, M. Qiao, Y. Wu, and M. Wang · 2025
Closest in time.
Hybrid latent reasoning via reinforcement learning, 2025
Z. Yue, B. Jin, H. Zeng, H. Zhuang, Z. Qin, J. Yoon, L. Shang, J. Han, and D. Wang · 2025
Closest in time.
Reasoning by superposition: A theoretical perspective on chain of continuous thought, 2025
H. Zhu, S. Hao, Z. Hu, J. Jiao, S. Russell, and Y. Tian · 2025
Closest in time.