Fetching the paper…
Reading the bibliography…
While state-of-the-art LLMs have demonstrated great promise of using long Chains-of-Thought (CoT) to boost reasoning, scaling it up to more challenging problems at test-time is fundamentally limited by suboptimal memory usage -- intermediate computations accumulate indefinitely in context even when no longer needed for future thoughts.
Word problems requiring exponential time (preliminary report)
L. J. Stockmeyer and A. R. Meyer · 1973
Earlier work this paper cites.
Equational logic as a programming language
M. J. O’Donnell · 1985
Earlier work this paper cites.
Automated reasoning introduction and applications
L. Wos, R. Overbeek, E. Lusk, and J. Boyle · 1992
Earlier work this paper cites.
Hybrid algorithms for the constraint satisfaction problem
P. Prosser · 1993
Earlier work this paper cites.
Generating hard satisfiability problems
B. Selman, D. G. Mitchell, and H. J. Levesque · 1996
Earlier work this paper cites.
Term rewriting and all that
F. Baader and T. Nipkow · 1998
Earlier work this paper cites.
On the complexity of k-sat
R. Impagliazzo and R. Paturi · 2001
Earlier work this paper cites.
Computational complexity: a modern approach
S. Arora and B. Barak · 2009
Earlier work this paper cites.
Language modeling with gated convolutional networks
Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier · 2017
Earlier work this paper cites.
Searching for activation functions
P. Ramachandran, B. Zoph, and Q. V. Le · 2017
Earlier work this paper cites.
Longformer: The long-document transformer
I. Beltagy, M. E. Peters, and A. Cohan · 2020
Earlier work this paper cites.
Rethinking attention with performers
K. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins, J. Davis, A. Mohiuddin, L. Kaiser, et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
Reformer: The efficient transformer
N. Kitaev, 𝖫 \mathsf{L} . Kaiser, and A. Levskaya · 2020
Earlier work this paper cites.
Glu variants improve transformer
N. Shazeer · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
M. Zaheer, G. Guruganesh, K. A. Dubey, J. Ainslie, C. Alberti, S. Ontanon, P. Pham, A. Ravula, Q. Wang, L. Yang, et al · 2020
Earlier work this paper cites.
Efficiently modeling long sequences with structured state spaces
A. Gu, K. Goel, and C. Ré · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
M. Nye, A. J. Andreassen, G. Gur-Ari, H. Michalewski, J. Austin, D. Bieber, D. Dohan, A. Lewkowycz, M. Bosma, D. Luan, et al · 2021
Earlier work this paper cites.
Attention is turing-complete
J. Pérez, P. Barceló, and J. Marinkovic · 2021
Earlier work this paper cites.
Thinking like transformers
G. Weiss, Y. Goldberg, and E. Yahav · 2021
Cited alongside, same era.
W. Chen, X. Ma, X. Wang, and W. W. Cohen · 2022
Cited alongside, same era.
Compositional semantic parsing with large language models
A. Drozdov, N. Schärli, E. Akyürek, N. Scales, X. Song, X. Chen, O. Bousquet, and D. Zhou · 2022
Cited alongside, same era.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Cited alongside, same era.
Decomposed prompting: A modular approach for solving complex tasks
T. Khot, H. Trivedi, M. Finlayson, Y. Fu, K. Richardson, P. Clark, and A. Sabharwal · 2022
Faith and fate: Limits of transformers on compositionality
N. Dziri, X. Lu, M. Sclar, X. L. Li, L. Jiang, B. Y. Lin, S. Welleck, P. West, C. Bhagavatula, R. Le Bras, et al · 2024
Later among the works it cites.
Towards revealing the mystery behind chain of thought: a theoretical perspective
G. Feng, B. Zhang, Y. Gu, H. Ye, D. He, and L. Wang · 2024
Later among the works it cites.
Lazyllm: Dynamic token pruning for efficient long context llm inference
Q. Fu, M. Cho, T. Merth, S. Mehta, M. Rastegari, and M. Najibi · 2024
Later among the works it cites.
Memory makes computation universal, remember?
E. Garrison · 2024
Later among the works it cites.
Lost in the middle: How language models use long contexts
N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learned token pruning for transformers
S. Kim, S. Shen, D. Thorsley, A. Gholami, W. Kwon, J. Hassoun, and K. Keutzer · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa · 2022
Cited alongside, same era.
Saturated transformers are constant-depth threshold circuits
W. Merrill, A. Sabharwal, and N. A. Smith · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Cited alongside, same era.
Star: Bootstrapping reasoning with reasoning
E. Zelikman, Y. Wu, J. Mu, and N. Goodman · 2022
Cited alongside, same era.
Least-to-most prompting enables complex reasoning in large language models
D. Zhou, N. Schärli, L. Hou, J. Wei, N. Scales, X. Wang, D. Schuurmans, C. Cui, O. Bousquet, Q. Le, et al · 2022
Cited alongside, same era.
Retrieval-augmented generation for large language models: A survey
Y. Gao, Y. Xiong, X. Gao, K. Jia, J. Pan, Y. Bi, Y. Dai, J. Sun, and H. Wang · 2023
Cited alongside, same era.
Self-refine: Iterative refinement with self-feedback
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, et al · 2024
Later among the works it cites.
Dynamic memory compression: Retrofitting llms for accelerated inference
P. Nawrot, A. 𝖫 \mathsf{L} ańcucki, M. Chochowski, D. Tarjan, and E. M. Ponti · 2024
Later among the works it cites.
On the representational capacity of neural language models with chain-of-thought reasoning
F. Nowak, A. Svete, A. Butoi, and R. Cotterell · 2024
Later among the works it cites.
Learning to reason with llms, September 2024
OpenAI · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
C. Snell, J. Lee, K. Xu, and A. Kumar · 2024
Later among the works it cites.
What formal languages can transformers express? a survey
L. Strobl, W. Merrill, G. Weiss, D. Chiang, and D. Angluin · 2024
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu · 2024
Later among the works it cites.
Meta-prompting: Enhancing language models with task-agnostic scaffolding
M. Suzgun and A. T. Kalai · 2024
Later among the works it cites.
Augmenting language models with long-term memory
W. Wang, L. Dong, H. Cheng, X. Liu, X. Yan, J. Gao, and F. Wei · 2024
Later among the works it cites.
Counting like transformers: Compiling temporal counting logic into softmax transformers
A. Yang and D. Chiang · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao, and K. Narasimhan · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
N. Muennighoff, Z. Yang, W. Shi, X. L. Li, L. Fei-Fei, H. Hajishirzi, L. Zettlemoyer, P. Liang, E. Candès, and T. Hashimoto · 2025
Closest in time.