Fetching the paper…
Reading the bibliography…
Large reasoning models (LRMs) have exhibited remarkable reasoning capabilities through inference-time scaling, but this progress has also introduced considerable redundancy and inefficiency into their reasoning processes, resulting in substantial computational waste.
Race: Large-scale reading comprehension dataset from examinations
Lai, G.; Xie, Q.; Liu, H.; Yang, Y.; and Hovy, E. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
Rajbhandari, S.; Rasley, J.; Ruwase, O.; and He, Y. 2020 · 2020
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D.; Burns, C.; Kadavath, S.; Arora, A.; Basart, S.; Tang, E.; Song, D.; and Steinhardt, J. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W.; et al. 2022 · 2022
Earlier work this paper cites.
Solving quantitative reasoning problems with language models
Lewkowycz, A.; Andreassen, A.; Dohan, D.; Dyer, E.; Michalewski, H.; Ramasesh, V.; Slone, A.; Anil, C.; Schlag, I.; Gutman-Solo, T.; et al. 2022 · 2022
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods, 2022
Lin, S.; Hilton, J.; and Evans, O. 2021 · 2022
Earlier work this paper cites.
Do not think that much for 2+ 3=? on the overthinking of o1-like llms
Chen, X.; Xu, J.; Liang, T.; He, Z.; Pang, J.; Yu, D.; Song, L.; Liu, Q.; Zhou, M.; Zhang, Z.; et al. 2024 · 2024
Earlier work this paper cites.
Efficiently Serving LLM Reasoning Programs with Certaindex
Fu, Y.; Chen, J.; Zhu, S.; Fu, Z.; Dai, Z.; Qiao, A.; and Zhang, H. 2024 · 2024
Earlier work this paper cites.
Token-budget-aware llm reasoning
Han, T.; Wang, Z.; Fang, C.; Zhao, S.; Ma, S.; and Chen, Z. 2024 · 2024
Earlier work this paper cites.
He, C.; Luo, R.; Bai, Y.; Hu, S.; Thai, Z. L.; Shen, J.; Hu, J.; Han, X.; Huang, Y.; Zhang, Y.; et al. 2024 · 2024
Earlier work this paper cites.
Jaech, A.; Kalai, A.; Lerer, A.; Richardson, A.; El-Kishky, A.; Low, A.; Helyar, A.; Madry, A.; Beutel, A.; Carney, A.; et al. 2024 · 2024
Earlier work this paper cites.
Livecodebench: Holistic and contamination free evaluation of large language models for code
Jain, N.; Han, K.; Gu, A.; Li, W.-D.; Yan, F.; Zhang, T.; Wang, S.; Solar-Lezama, A.; Sen, K.; and Stoica, I. 2024 · 2024
Earlier work this paper cites.
Improve mathematical reasoning in language models by automated process supervision
Luo, L.; Liu, Y.; Liu, R.; Phatale, S.; Guo, M.; Lara, H.; Li, Y.; Shu, L.; Zhu, Y.; Meng, L.; et al. 2024 · 2024
Cited alongside, same era.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z.; Wang, P.; Zhu, Q.; Xu, R.; Song, J.; Bi, X.; Zhang, H.; Zhang, M.; Li, Y.; Wu, Y.; et al. 2024 · 2024
Cited alongside, same era.
Aime problem set 1983-2024, 2023
Veeraboina, H. 2023 · 2024
Cited alongside, same era.
Qwen2. 5-math technical report: Toward mathematical expert model via self-improvement
Yang, A.; Zhang, B.; Hui, B.; Gao, B.; Yu, B.; Li, C.; Liu, D.; Tu, J.; Zhou, J.; Lin, J.; et al. 2024 · 2024
Cited alongside, same era.
L1: Controlling how long a reasoning model thinks with reinforcement learning
Understanding r1-zero-like training: A critical perspective
Liu, Z.; Chen, C.; Li, W.; Qi, P.; Pang, T.; Du, C.; Lee, W. S.; and Lin, M. 2025 · 2025
Closest in time.
AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning
Lou, C.; Sun, Z.; Liang, X.; Qu, M.; Shen, W.; Wang, W.; Li, Y.; Yang, Q.; and Wu, S. 2025 · 2025
Closest in time.
A survey of efficient reasoning for large reasoning models: Language, multimodality, and beyond
Qu, X.; Li, Y.; Su, Z.; Sun, W.; Yan, J.; Liu, D.; Cui, G.; Liu, D.; Liang, S.; He, J.; et al. 2025 · 2025
Closest in time.
Stop overthinking: A survey on efficient reasoning for large language models
Sui, Y.; Chuang, Y.-N.; Wang, G.; Zhang, J.; Zhang, T.; Yuan, J.; Liu, H.; Wen, A.; Zhong, S.; Chen, H.; et al. 2025 · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aggarwal, P.; and Welleck, S. 2025 · 2025
Cited alongside, same era.
Training Language Models to Reason Efficiently
Arora, D.; and Zanette, A. 2025 · 2025
Cited alongside, same era.
The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks
Cuadron, A.; Li, D.; Ma, W.; Wang, X.; Wang, Y.; Zhuang, S.; Liu, S.; Schroeder, L. G.; Xia, T.; Mao, H.; et al. 2025 · 2025
Cited alongside, same era.
Thinkless: Llm learns when to think
Fang, G.; Ma, X.; and Wang, X. 2025 · 2025
Cited alongside, same era.
Efficient reasoning models: A survey
Feng, S.; Fang, G.; Ma, X.; and Wang, X. 2025 · 2025
Cited alongside, same era.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D.; Yang, D.; Zhang, H.; Song, J.; Zhang, R.; Xu, R.; Zhu, Q.; Ma, S.; Wang, P.; Bi, X.; et al. 2025 · 2025
Cited alongside, same era.
Thinkprune: Pruning long chain-of-thought of llms via reinforcement learning
Hou, B.; Zhang, Y.; Ji, J.; Liu, Y.; Qian, K.; Andreas, J.; and Chang, S. 2025 · 2025
Cited alongside, same era.
Overthink: Slowdown attacks on reasoning llms. arXiv e-prints, pages arXiv–2502
Kumar, A.; Roh, J.; Naseh, A.; Karpinska, M.; Iyyer, M.; Houmansadr, A.; and Bagdasarian, E. 2025 · 2025
Cited alongside, same era.
Wu, J.; Zhu, J.; and Liu, Y. 2025 · 2025
Closest in time.
Tokenskip: Controllable chain-of-thought compression in llms
Xia, H.; Li, Y.; Leong, C. T.; Wang, W.; and Li, W. 2025 · 2025
Closest in time.
Chain of draft: Thinking faster by writing less
Xu, S.; Xie, W.; Zhao, L.; and He, P. 2025 · 2025
Closest in time.
Inftythink: Breaking the length limits of long-context reasoning in large language models
Yan, Y.; Shen, Y.; Liu, Y.; Jiang, J.; Zhang, M.; Shao, J.; and Zhuang, Y. 2025 · 2025
Closest in time.
LIMO: Less is More for Reasoning
Ye, Y.; Huang, Z.; Xiao, Y.; Chern, E.; Xia, S.; and Liu, P. 2025 · 2025
Closest in time.
Dapo: An open-source llm reinforcement learning system at scale
Yu, Q.; Zhang, Z.; Zhu, R.; Yuan, Y.; Zuo, X.; Yue, Y.; Dai, W.; Fan, T.; Liu, G.; Liu, L.; et al. 2025 · 2025
Closest in time.
Zhang, J.; and Zuo, C. 2025 · 2025
Closest in time.
The surprising effectiveness of negative reinforcement in LLM reasoning
Zhu, X.; Xia, M.; Wei, Z.; Chen, W.-L.; Chen, D.; and Meng, Y. 2025 · 2025
Closest in time.