Fetching the paper…
Reading the bibliography…
Recent work has demonstrated the remarkable potential of Large Language Models (LLMs) in test-time scaling.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Henighan, et al. 2020 · 2001
Earlier work this paper cites.
Positional encoding to control output sequence length
Sho Takase and Naoaki Okazaki. 2019 · 2019
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, et al. 2021 · 2021
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Schuurmans, et al. 2022 · 2022
Earlier work this paper cites.
Reasoning with language model is planning with world model
Shibo Hao, Yi Gu, Ma, et al. 2023 · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, et al. 2023 · 2023
Earlier work this paper cites.
Hunter Lightman, Vineet Kosaraju, Yura Burda, et al. 2023 · 2023
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, et al. 2023 · 2023
Earlier work this paper cites.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Zhao, et al. 2023 · 2023
Earlier work this paper cites.
Large language monkeys: Scaling inference compute with repeated sampling
Bradley Brown, Jordan Juravsky, Ehrlich, and othe. 2024 · 2024
Earlier work this paper cites.
Precise length control in large language models
Bradley Butcher, Michael O’Keefe, and James Titchener. 2024 · 2024
Earlier work this paper cites.
Learning how hard to think: Input-adaptive allocation of lm computation
Mehul Damani, Idan Shenfeld, Andi Peng, Andreea Bobu, and Jacob Andreas. 2024 · 2024
Cited alongside, same era.
Improving factuality and reasoning in language models through multiagent debate
Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch. 2024 · 2024
Cited alongside, same era.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, et al. 2024 · 2024
Cited alongside, same era.
Token-budget-aware llm reasoning
Tingxu Han, Zhenting Wang, Chunrong Fang, Shiyu Zhao, Shiqing Ma, and Zhenyu Chen. 2024 · 2024
Cited alongside, same era.
When can LLMs actually correct their own mistakes? a critical survey of self-correction of LLMs
Ryo Kamoi, Yusen Zhang, Nan Zhang, Jiawei Han, and Rui Zhang. 2024 · 2024
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. 2024 · 2024
Later among the works it cites.
Drivecot: Integrating chain-of-thought reasoning with end-to-end driving
Tianqi Wang, Enze Xie, Ruihang Chu, Zhenguo Li, and Ping Luo. 2024 · 2024
Later among the works it cites.
Following length constraints in instructions
Weizhe Yuan, Ilia Kulikov, Ping Yu, Kyunghyun Cho, Sainbayar Sukhbaatar, Jason Weston, and Jing Xu. 2024 · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
DeepSeek-AI, Daya Guo, Dejian Yang, et al. 2025 · 2025
Closest in time.
Advancing language model reasoning through reinforcement learning and inference scaling
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Escape sky-high cost: Early-stopping self-consistency for multi-step reasoning
Yiwei Li, Peiwen Yuan, Shaoxiong Feng, et al. 2024 · 2024
Cited alongside, same era.
Encouraging divergent thinking in large language models through multi-agent debate
Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Shuming Shi, and Zhaopeng Tu. 2024 · 2024
Cited alongside, same era.
Let’s verify step by step
Hunter Lightman, Vineet Kosaraju, Yuri Burda, et al. 2024 · 2024
Cited alongside, same era.
Adaptive inference-time compute: Llms can predict if they can do better, even mid-generation
Rohin Manvi, Anikait Singh, and Stefano Ermon. 2024 · 2024
Cited alongside, same era.
Llama 3.2 model card
Meta. 2024 · 2024
Cited alongside, same era.
Learning to reason with llms
OpenAI. 2024 · 2024
Cited alongside, same era.
Phi-3 technical report: A highly capable language model locally on your phone
Marah Abdin, Jyoti Aneja, Hany Awadalla, et al. 2024a
Cited in the paper.
Zhenyu Hou, Xin Lv, Rui Lu, Zhang, et al. 2025 · 2025
Closest in time.
Acpbench: Reasoning about action, change, and planning
Harsha Kokel, Michael Katz, Kavitha Srinivas, and Shirin Sohrabi. 2025 · 2025
Closest in time.
Kuang-Huei Lee, Ian Fischer, Wu, et al. 2025 · 2025
Closest in time.
Concise thoughts: Impact of output length on llm reasoning and cost
Sania Nayab, Giulio Rossolini, Marco Simoni, et al. 2025 · 2025
Closest in time.
Sky-t1: Fully open-source reasoning model with o1-preview performance in $450 budget
NovaSky Team. 2025 · 2025
Closest in time.
Make every penny count: Difficulty-adaptive self-consistency for cost-efficient reasoning
Xinglin Wang, Shaoxiong Feng, Yiwei Li, et al. 2025 · 2025
Closest in time.
Chain of draft: Thinking faster by writing less
Silei Xu, Wenhao Xie, Lingxiao Zhao, and Pengcheng He. 2025 · 2025
Closest in time.