Fetching the paper…
Reading the bibliography…
Test-time scaling (TTS), which involves dynamic allocation of compute during inference, offers a promising way to improve reasoning in large language models.
The asymptotic theory of extreme order statistics
J. Galambos · 1977
Earlier work this paper cites.
Diverse beam search: Decoding diverse solutions from neural sequence models
A. K. Vijayakumar, M. Cogswell, R. R. Selvaraju, Q. Sun, S. Lee, D. Crandall, and D. Batra · 2016
Earlier work this paper cites.
J. Liu, A. Cohen, R. Pasunuru, Y. Choi, H. Hajishirzi, and A. Celikyilmaz · 2023
Earlier work this paper cites.
Planning with large language models for code generation
S. Zhang, Z. Chen, Y. Shen, M. Ding, J. B. Tenenbaum, and C. Gan · 2023
Earlier work this paper cites.
Large language monkeys: Scaling inference compute with repeated sampling
B. Brown, J. Juravsky, R. Ehrlich, R. Clark, Q. V. Le, C. Ré, and A. Mirhoseini · 2024
Earlier work this paper cites.
Inference-aware fine-tuning for best-of-n sampling in large language models
Y. Chow, G. Tennenholtz, I. Gur, V. Zhuang, B. Dai, S. Thiagarajan, C. Boutilier, R. Agarwal, A. Kumar, and A. Faust · 2024
Earlier work this paper cites.
Stream of search (sos): Learning to search in language
K. Gandhi, D. Lee, G. Grand, M. Liu, W. Cheng, A. Sharma, and N. D. Goodman · 2024
Earlier work this paper cites.
Olympicarena medal ranks: Who is the most intelligent ai so far?
Z. Huang, Z. Wang, S. Xia, and P. Liu · 2024
Earlier work this paper cites.
A simple model of inference scaling laws
N. Levi · 2024
Earlier work this paper cites.
A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, et al · 2024
Earlier work this paper cites.
Gpqa: A graduate-level google-proof q&a benchmark
D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman · 2024
Cited alongside, same era.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
C. Snell, J. Lee, K. Xu, and A. Kumar · 2024
Cited alongside, same era.
From decoding to meta-generation: Inference-time algorithms for large language models
S. Welleck, A. Bertsch, M. Finlayson, H. Schoelkopf, A. Xie, G. Neubig, I. Kulikov, and Z. Harchaoui · 2024
Cited alongside, same era.
An empirical analysis of compute-optimal inference for problem-solving with language models
Y. Wu, Z. Sun, S. Li, S. Welleck, and Y. Yang · 2024
Cited alongside, same era.
Monte carlo tree search boosts reasoning via iterative preference learning
Advancing language model reasoning through reinforcement learning and inference scaling, 2025
Z. Hou, X. Lv, R. Lu, J. Zhang, Y. Li, Z. Yao, J. Li, J. Tang, and Y. Dong · 2025
Closest in time.
K.-H. Lee, I. Fischer, Y.-H. Wu, D. Marwood, S. Baluja, D. Schuurmans, and X. Chen · 2025
Closest in time.
Aime problems and solutions
MAA Committee · 2025
Closest in time.
N. Muennighoff, Z. Yang, W. Shi, X. L. Li, L. Fei-Fei, H. Hajishirzi, L. Zettlemoyer, P. Liang, E. Candès, and T. Hashimoto · 2025
Closest in time.
Learning to reason with llms
OpenAI · 2025
Closest in time.
Coat: Chain-of-associated-thoughts framework for enhancing large language models reasoning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Xie, A. Goyal, W. Zheng, M.-Y. Kan, T. P. Lillicrap, K. Kawaguchi, and M. Shieh · 2024
Cited alongside, same era.
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu · 2024
Cited alongside, same era.
Phi-4-reasoning technical report
M. Abdin, S. Agarwal, A. Awadallah, V. Balachandran, H. Behl, L. Chen, G. de Rosa, S. Gunasekar, M. Javaheripi, N. Joshi, et al · 2025
Cited alongside, same era.
L1: Controlling how long a reasoning model thinks with reinforcement learning
P. Aggarwal and S. Welleck · 2025
Cited alongside, same era.
Thinking machines: A survey of llm based reasoning strategies
D. Bandyopadhyay, S. Bhattacharjee, and A. Ekbal · 2025
Cited alongside, same era.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Cited alongside, same era.
S*: Test time scaling for code generation
D. Li, S. Cao, C. Cao, X. Li, S. Tan, K. Keutzer, J. Xing, J. E. Gonzalez, and I. Stoica
Cited in the paper.
Learning to reason from feedback at test-time
Y. Li, M. Lyu, and L. Wang
Cited in the paper.
J. Pan, S. Deng, and S. Huang · 2025
Closest in time.
H. Peng, Y. Qi, X. Wang, Z. Yao, B. Xu, L. Hou, and J. Li · 2025
Closest in time.
Qwq-32b: Embracing the power of reinforcement learning, March 2025
Q. Team · 2025
Closest in time.
What, how, where, and how well? a survey on test-time scaling in large language models
Q. Zhang, F. Lyu, Z. Sun, L. Wang, W. Zhang, Z. Guo, Y. Wang, I. King, X. Liu, and C. Ma · 2025
Closest in time.