Fetching the paper…
Reading the bibliography…
Test-Time Scaling (TTS) methods for enhancing Large Language Model (LLM) reasoning often incur substantial computational costs, primarily due to extensive reliance on external Process Reward Models (PRMs) or sampling methods like Best-of-N (BoN).
Sentence-bert: Sentence embeddings using siamese bert-networks
N. Reimers and I. Gurevych · 1908
Earlier work this paper cites.
SGDR: stochastic gradient descent with warm restarts
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Trl: Transformer reinforcement learning
L. von Werra, Y. Belkada, L. Tunstall, E. Beeching, T. Thrush, N. Lambert, S. Huang, K. Rasul, and Q. Gallouédec · 2020
Earlier work this paper cites.
Measuring mathematical problem solving with the MATH dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Earlier work this paper cites.
Will we run out of data? limits of llm scaling based on human-generated data
P. Villalobos, A. Ho, J. Sevilla, T. Besiroglu, L. Heim, and M. Hobbhahn · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
Amc 2023, 2024b
AI-MO · 2023
Earlier work this paper cites.
Flashattention-2: Faster attention with better parallelism and work partitioning
T. Dao · 2023
Earlier work this paper cites.
Fewer is more: Boosting llm reasoning with reinforced context pruning
X. Huang, L. L. Zhang, K.-T. Cheng, F. Yang, and M. Yang · 2023
Earlier work this paper cites.
Let’s verify step by step
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe · 2023
Earlier work this paper cites.
Math-shepherd: Verify and reinforce llms step-by-step without human annotations
P. Wang, L. Li, Z. Shao, R. Xu, D. Dai, Y. Li, D. Chen, Y. Wu, and Z. Sui · 2023
Earlier work this paper cites.
Self-evaluation guided beam search for reasoning
Y. Xie, K. Kawaguchi, Y. Zhao, J. X. Zhao, M.-Y. Kan, J. He, and M. Xie · 2023
Earlier work this paper cites.
Aime 2024, 2024a
AI-MO · 2024
Cited alongside, same era.
Large language monkeys: Scaling inference compute with repeated sampling
B. Brown, J. Juravsky, R. Ehrlich, R. Clark, Q. V. Le, C. Ré, and A. Mirhoseini · 2024
Cited alongside, same era.
Reft: Reasoning with reinforced fine-tuning
T. Q. Luong, X. Zhang, Z. Jie, P. Sun, X. Jin, and H. Li · 2024
Cited alongside, same era.
Learning to reason with llms, 2024
OpenAI · 2024
Cited alongside, same era.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. Li, Y. Wu, et al · 2024
Cited alongside, same era.
Test-time computing: from system-1 thinking to system-2 thinking
Y. Ji, J. Li, H. Ye, K. Wu, J. Xu, L. Mo, and M. Zhang · 2025
Closest in time.
DeepScaleR: Surpassing O1-preview with a 1.5B model by scaling RL, 2025
M. Luo, S. Tan, J. Wong, X. Shi, W. Y. Tang, M. Roongta, C. Cai, J. Luo, L. E. Li, R. A. Popa, and I. Stoica · 2025
Closest in time.
N. Muennighoff, Z. Yang, W. Shi, X. L. Li, L. Fei-Fei, H. Hajishirzi, L. Zettlemoyer, P. Liang, E. Candès, and T. Hashimoto · 2025
Closest in time.
Premise-augmented reasoning chains improve error identification in math reasoning with llms
S. Mukherjee, A. Chinta, T. Kim, T. A. Sharma, and D. Hakkani-Tür · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Snell, J. Lee, K. Xu, and A. Kumar · 2024
Cited alongside, same era.
Dynamic self-consistency: Leveraging reasoning paths for efficient llm sampling
G. Wan, Y. Wu, J. Chen, and S. Li · 2024
Cited alongside, same era.
Openr: An open source framework for advanced reasoning with large language models
J. Wang, M. Fang, Z. Wan, M. Wen, J. Zhu, A. Liu, Z. Gong, Y. Song, L. Chen, L. M. Ni, et al · 2024
Cited alongside, same era.
Monte carlo tree search boosts reasoning via iterative preference learning
Y. Xie, A. Goyal, W. Zheng, M.-Y. Kan, T. P. Lillicrap, K. Kawaguchi, and M. Shieh · 2024
Cited alongside, same era.
An implementation of generative prm, 2024
W. Xiong, H. Zhang, N. Jiang, and T. Zhang · 2024
Cited alongside, same era.
Processbench: Identifying process errors in mathematical reasoning
C. Zheng, Z. Zhang, B. Zhang, R. Lin, K. Lu, B. Yu, D. Liu, J. Zhou, and J. Lin · 2024
Cited alongside, same era.
A survey on efficient inference for large language models
Z. Zhou, X. Ning, K. Hong, T. Fu, J. Xu, S. Li, Y. Lou, L. Wang, Z. Yuan, X. Li, et al · 2024
Cited alongside, same era.
A. Razghandi, S. M. H. Hosseini, and M. S. Baghshah · 2025
Closest in time.
Confidence improves self-consistency in llms
A. Taubenfeld, T. Sheffer, E. Ofek, A. Feder, A. Goldstein, Z. Gekhman, and G. Yona · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
K. Team, A. Du, B. Gao, B. Xing, C. Jiang, C. Chen, C. Li, C. Xiao, C. Du, C. Liao, et al · 2025
Closest in time.
When more is less: Understanding chain-of-thought length in llms
Y. Wu, Y. Wang, T. Du, S. Jegelka, and Y. Wang · 2025
Closest in time.
Limo: Less is more for reasoning, 2025
Y. Ye, Z. Huang, Y. Xiao, E. Chern, S. Xia, and P. Liu · 2025
Closest in time.
Aime 2025, 2025
yentinglin · 2025
Closest in time.
Dapo: An open-source llm reinforcement learning system at scale
Q. Yu, Z. Zhang, R. Zhu, Y. Yuan, X. Zuo, Y. Yue, T. Fan, G. Liu, L. Liu, X. Liu, et al · 2025
Closest in time.
Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild
W. Zeng, Y. Huang, Q. Liu, W. Liu, K. He, Z. Ma, and J. He · 2025
Closest in time.
The lessons of developing process reward models in mathematical reasoning
Z. Zhang, C. Zheng, Y. Wu, B. Zhang, R. Lin, B. Yu, D. Liu, J. Zhou, and J. Lin · 2025
Closest in time.
A comprehensive survey of reward models: Taxonomy, applications, challenges, and future
J. Zhong, W. Shen, Y. Li, S. Gao, H. Lu, Y. Chen, Y. Zhang, W. Zhou, J. Gu, and L. Zou · 2025
Closest in time.
amc23 dataset
zwhe99 · 2025
Closest in time.