Fetching the paper…
Reading the bibliography…
Chain of Thought (CoT) reasoning enhances language models' performance but often leads to inefficient "overthinking" on simple problems.
Alpacafarm: A simulation framework for methods that learn from human feedback
Y. Dubois, C. X. Li, R. Taori, T. Zhang, I. Gulrajani, J. Ba, C. Guestrin, P. S. Liang, and T. B. Hashimoto · 2023
Earlier work this paper cites.
Llm-blender: Ensembling large language models with pairwise ranking and generative fusion
D. Jiang, X. Ren, and B. Y. Lin · 2023
Earlier work this paper cites.
Alpacaeval: An automatic evaluator of instruction-following models
X. Li, T. Zhang, Y. Dubois, R. Taori, I. Gulrajani, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Earlier work this paper cites.
A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, et al · 2024
Earlier work this paper cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. Li, Y. Wu, et al · 2024
Earlier work this paper cites.
Hybridflow: A flexible and efficient rlhf framework
G. Sheng, C. Zhang, Z. Ye, X. Wu, W. Zhang, R. Zhang, Y. Peng, H. Lin, and C. Wu · 2024
Cited alongside, same era.
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, et al · 2024
Cited alongside, same era.
L1: Controlling how long a reasoning model thinks with reinforcement learning
P. Aggarwal and S. Welleck · 2025
Cited alongside, same era.
Training language models to reason efficiently, 2025
D. Arora and A. Zanette · 2025
Cited alongside, same era.
Adaptive group policy optimization: Towards stable training and token-efficient reasoning
K. Y. Chen Li, Nazhou Liu · 2025
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
Pairwise rm: Perform best-of-n sampling with knockout tournament
Y. Liu, Z. Yao, R. Min, Y. Cao, L. Hou, and J. Li · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
K. Team, A. Du, B. Gao, B. Xing, C. Jiang, C. Chen, C. Li, C. Xiao, C. Du, C. Liao, et al · 2025
Closest in time.
Dapo: An open-source llm reinforcement learning system at scale
Q. Yu, Z. Zhang, R. Zhu, Y. Yuan, X. Zuo, Y. Yue, T. Fan, G. Liu, L. Liu, X. Liu, et al · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
O1-pruner: Length-harmonizing fine-tuning for o1-like reasoning pruning
H. Luo, L. Shen, H. He, Y. Wang, S. Liu, W. Li, N. Tan, X. Cao, and D. Tao
Cited in the paper.
Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl, 2025b
M. Luo, S. Tan, J. Wong, X. Shi, W. Y. Tang, M. Roongta, C. Cai, J. Luo, T. Zhang, L. E. Li, R. A. Popa, and I. Stoica
Cited in the paper.