Fetching the paper…
Reading the bibliography…
Enhancing the reasoning capabilities of Large Language Models (LLMs) with efficiency and scalability remains a fundamental challenge in artificial intelligence research.
Training language models to follow instructions with human feedback
L. Ouyang et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei et al · 2022
Earlier work this paper cites.
OpenAI · 2023
Earlier work this paper cites.
Reflexion: Language agents with verbal reinforcement learning
N. Shinn et al · 2023
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
X. Wang et al · 2023
Earlier work this paper cites.
Tree of thoughts: Deliberate problem solving with large language models
S. Yao et al · 2023
Earlier work this paper cites.
Think twice: An inference-time approach to improve reasoning in llms
D. Zhou et al · 2023
Earlier work this paper cites.
Numinamath
Jia LI, Edward Beeching, Lewis Tunstall, Ben Lipkin, Roman Soletskyi, Shengyi Costa Huang, Kashif Rasul, Longhui Yu, Albert Jiang, Ziju Shen, Zihan Qin, Bin Dong, Li Zhou, Yann Fleureau, Guillaume Lample, and Stanislas Polu · 2024
Earlier work this paper cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Y Wu, et al · 2024
Earlier work this paper cites.
Kimi-k1.5: Reinforcement learning enhanced llm reasoning
Kimi AI · 2025
Cited alongside, same era.
Claude 3: Advanced reasoning and alignment
Anthropic · 2025
Cited alongside, same era.
Datawhale-r1: A chinese tutorial-level reproduction of deepseek-r1 zero, 2025
Datawhale Community and SLRLab · 2025
Cited alongside, same era.
Process reinforcement through implicit rewards
Ganqu Cui, Lifan Yuan, Zefan Wang, Hanbin Wang, Wendi Li, Bingxiang He, Yuchen Fan, Tianyu Yu, Qixin Xu, Weize Chen, et al · 2025
Cited alongside, same era.
Math-verify, 2025a
Hugging Face · 2025
Cited alongside, same era.
Kw-r1: A simple implementation of the grpo algorithm
Jiaqing Liang, Jinyi Han, Xinyi Wang, Zishang Jiang, Chengyuan Xiong, Boyu Zhu, Jie Shi, Weijia Li, Tingyun Li, and Yanghua Xiao · 2025
Tinyzero
Jiayi Pan, Junjie Zhang, Xingyao Wang, Lifan Yuan, Hao Peng, and Alane Suhr · 2025
Closest in time.
Group relative policy optimization for language model training
D. Silver et al · 2025
Closest in time.
Deepseek-r1: Reinforcement learning for advanced reasoning
DeepSeek Team · 2025
Closest in time.
Open Thoughts
OpenThoughts Team · 2025
Closest in time.
Think twice: Enhancing llm reasoning by scaling multi-round test-time thinking, 2025
Xiaoyu Tian, Sitong Zhao, Haotian Wang, Shuaiting Chen, Yunjie Ji, Yiping Peng, Han Zhao, and Xiangang Li · 2025
Closest in time.
Limo: Less is more for reasoning
Yixin Ye, Zhen Huang, Yang Xiao, Ethan Chern, Shijie Xia, and Pengfei Liu · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Understanding r1-zero-like training: A critical perspective
Zichen Liu, Changyu Chen, Wenjun Li, Penghui Qi, Tianyu Pang, Chao Du, Wee Sun Lee, and Min Lin · 2025
Cited alongside, same era.
Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl
Michael Luo, Sijun Tan, Justin Wong, Xiaoxiang Shi, William Y. Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica · 2025
Cited alongside, same era.
Open-r1: A fully open reproduction of deepseek-r1, 2025b
Hugging Face
Cited in the paper.
Weihao Zeng, Yuzhen Huang, Qian Liu, Wei Liu, Keqing He, Zejun Ma, and Junxian He · 2025
Closest in time.
1.4 million open-source distilled reasoning dataset to empower large language model training, 2025
Han Zhao, Haotian Wang, Yiping Peng, Sitong Zhao, Xiaoyu Tian, Shuaiting Chen, Yunjie Ji, and Xiangang Li · 2025
Closest in time.