Fetching the paper…
Reading the bibliography…
This paper introduces Group Sequence Policy Optimization (GSPO), our stable, efficient, and performant reinforcement learning algorithm for training large language models.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Click: Controllable text generation with sequence likelihood contrastive learning
Chujie Zheng, Pei Ke, Zheng Zhang, and Minlie Huang · 2023
Earlier work this paper cites.
Learning to reason with LLMs, 2024
OpenAI · 2024
Cited alongside, same era.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo · 2024
Cited alongside, same era.
Team Qwen
Cited in the paper.
Qwq-32b: Embracing the power of reinforcement learning, March 2025b
Team Qwen
Cited in the paper.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
DeepSeek-AI · 2025
Closest in time.
Minimax-m1: Scaling test-time compute efficiently with lightning attention
MiniMax · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…