Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) has significantly enhanced the reasoning capabilities of large language models (LLMs), but its reliance on expensive human-labeled data or complex reward models severely limits scalability.
Proximal policy optimization algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Earlier work this paper cites.
Large Language Models Can Self-Improve
Huang, J.; Gu, S. S.; Hou, L.; Wu, Y.; Wang, X.; Yu, H.; and Han, J. 2022 · 2022
Earlier work this paper cites.
Scaling laws for reward model overoptimization
Gao, L.; Schulman, J.; and Hilton, J. 2023 · 2023
Earlier work this paper cites.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafailov, R.; Sharma, A.; Mitchell, E.; Ermon, S.; Manning, C. D.; and Finn, C. 2023 · 2023
Earlier work this paper cites.
Large language model routing with benchmark datasets
Shnitzer, T.; Ou, A.; Silva, M.; Soule, K.; Sun, Y.; Solomon, J.; Thompson, N.; and Yurochkin, M. 2023 · 2023
Earlier work this paper cites.
Cai, Z.; Cao, M.; Chen, H.; Chen, K.; Chen, K.; and et al. 2024 · 2024
Earlier work this paper cites.
Dubey, A.; Jauhri, A.; Pandey, A.; et al. 2024 · 2024
Earlier work this paper cites.
El-Kishky, A.; Selsam, D.; Song, F.; and et al. 2024 · 2024
Earlier work this paper cites.
Smoothie: Label Free Language Model Routing
Guha, N.; Chen, M. F.; Chow, T.; Khare, I. S.; and R’e, C. 2024 · 2024
Cited alongside, same era.
He, C.; Luo, R.; Bai, Y.; Hu, S.; Thai, Z. L.; Shen, J.; Hu, J.; Han, X.; Huang, Y.; Zhang, Y.; et al. 2024 · 2024
Cited alongside, same era.
Open-Endedness is Essential for Artificial Superhuman Intelligence
Hughes, E.; Dennis, M.; Parker-Holder, J.; Behbahani, F. M. P.; Mavalankar, A.; Shi, Y.; Schaul, T.; and Rocktaschel, T. 2024 · 2024
Cited alongside, same era.
Preference Optimization for Reasoning with Pseudo Feedback
Jiao, F.; Guo, G.; Zhang, X.; Chen, N. F.; Joty, S.; and Wei, F. 2024 · 2024
Cited alongside, same era.
Self-Rewarding Language Models
Yuan, W.; Pang, R. Y.; Cho, K.; Sukhbaatar, S.; Xu, J.; and Weston, J. E. 2024 · 2024
Later among the works it cites.
ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Zeng, A.; Xu, B.; Wang, B.; Zhang, C.; Yin, D.; and et al. 2024 · 2024
Later among the works it cites.
Understanding R1-Zero-Like Training: A Critical Perspective
Liu, Z.-Y.; Chen, C.; Li, W.; Qi, P.; Pang, T.; Du, C.; Lee, W. S.; and Lin, M. 2025 · 2025
Closest in time.
Can Large Reasoning Models Self-Train?
Shafayat, S.; Tajwar, F.; Salakhutdinov, R.; Schneider, J.; and Zanette, A. 2025 · 2025
Closest in time.
Genius: A generalizable and purely unsupervised self-training framework for advanced reasoning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions
Li, J.; Beeching, E.; Tunstall, L.; Lipkin, B.; Soletskyi, R.; Huang, S.; Rasul, K.; Yu, L.; Jiang, A. Q.; Shen, Z.; et al. 2024 · 2024
Cited alongside, same era.
HybridFlow: A Flexible and Efficient RLHF Framework
Sheng, G.; Zhang, C.; Ye, Z.; Wu, X.; Zhang, W.; Zhang, R.; Peng, Y.; Lin, H.; and Wu, C. 2024 · 2024
Cited alongside, same era.
AI models collapse when trained on recursively generated data
Shumailov, I.; Shumaylov, Z.; Zhao, Y.; Papernot, N.; Anderson, R.; and Gal, Y. 2024 · 2024
Cited alongside, same era.
Yang, A.; Yang, B.; Zhang, B.; Hui, B.; Zheng, B.; and et al. 2024 · 2024
Cited alongside, same era.
Symbolic Mixture-of-Experts: Adaptive Skill-based Routing for Heterogeneous Reasoning
Chen, J. C.-Y.; Yun, S.; Stengel-Eskin, E.; Chen, T.; and Bansal, M. 2025a
Cited in the paper.
Symbolic mixture-of-experts: Adaptive skill-based routing for heterogeneous reasoning
Chen, J. C.-Y.; Yun, S.; Stengel-Eskin, E.; Chen, T.; and Bansal, M. 2025b
Cited in the paper.
Routerdc: Query-based router by dual contrastive learning for assembling large language models
Chen, S.; Jiang, W.; Lin, B.; Kwok, J.; and Zhang, Y. 2024a
Cited in the paper.
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
Chen, Z.; Deng, Y.; Yuan, H.; Ji, K.; and Gu, Q. 2024b
Cited in the paper.
Xu, F.; Yan, H.; Ma, C.; Zhao, H.; Sun, Q.; Cheng, K.; He, J.; Liu, J.; and Wu, Z. 2025 · 2025
Closest in time.
Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Yue, Y.; Chen, Z.; Lu, R.; Zhao, A.; Wang, Z.; Song, S.; and Huang, G. 2025 · 2025
Closest in time.
Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Zhang, Q.; Wu, H.; Zhang, C.; Zhao, P.; and Bian, Y. 2025 · 2025
Closest in time.
TTRL: Test-Time Reinforcement Learning
Zuo, Y.; Zhang, K.; Qu, S.; Sheng, L.; Zhu, X.; Qi, B.; Sun, Y.; Cui, G.; Ding, N.; and Zhou, B. 2025 · 2025
Closest in time.