Fetching the paper…
Reading the bibliography…
The reasoning capabilities of large language models (LLMs) have advanced rapidly, particularly following the release of DeepSeek R1, which has inspired a surge of research into data quality and reinforcement learning (RL) algorithms.
Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents
Edoardo Conti, Vashisht Madhavan, Felipe Petroski Such, Joel Lehman, Kenneth Stanley, and Jeff Clune · 2018
Earlier work this paper cites.
Diversity is all you need: Learning skills without a reward function
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine · 2018
Earlier work this paper cites.
Diversity-driven exploration strategy for deep reinforcement learning
Zhang-Wei Hong, Tzu-Yun Shann, Shih-Yang Su, Yi-Hsiang Chang, Tsu-Jui Fu, and Chun-Yi Lee · 2018
Earlier work this paper cites.
Understanding the impact of entropy on policy optimization
Zafarali Ahmed, NicolasLe Roux, Mohammad Norouzi, and Dale Schuurmans · 2019
Earlier work this paper cites.
Diversity-inducing policy gradient: Using maximum mean discrepancy to find a set of diverse policies
Muhammad A Masood and Finale Doshi-Velez · 2019
Earlier work this paper cites.
Learning novel policies for tasks
Yunbo Zhang, Wenhao Yu, and Greg Turk · 2019
Earlier work this paper cites.
Geoffrey Cideron, Thomas Pierrot, Nicolas Perrin, Karim Beguir, and Olivier Sigaud · 2020
Earlier work this paper cites.
Effective diversity in population based reinforcement learning
Jack Parker-Holder, Aldo Pacchiano, Krzysztof M Choromanski, and Stephen J Roberts · 2020
Earlier work this paper cites.
Non-local policy optimization via diversity-regularized collaborative exploration
Zhenghao Peng, Hao Sun, and Bolei Zhou · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He · 2020
Earlier work this paper cites.
Trl: Transformer reinforcement learning
Leandro von Werra, Younes Belkada, Lewis Tunstall, Edward Beeching, Tristan Thrush, Nathan Lambert, Shengyi Huang, Kashif Rasul, and Quentin Gallouédec · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
Multiple plans are better than one: Diverse stochastic planning
Mahsa Ghasemi, Evan Scope Crafts, Bo Zhao, and Ufuk Topcu · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Discovering diverse multi-agent strategic behavior via reward randomization
Zhenggang Tang, Chao Yu, Boyuan Chen, Huazhe Xu, Xiaolong Wang, Fei Fang, Simon Du, Yu Wang, and Yi Wu · 2021
Earlier work this paper cites.
Discovering diverse nearly optimal policies with successor features
Tom Zahavy, Brendan O’Donoghue, Andre Barreto, Volodymyr Mnih, Sebastian Flennerhag, and Satinder Singh · 2021
Earlier work this paper cites.
Diversity policy gradient for sample efficient quality-diversity optimization
Thomas Pierrot, Valentin Macé, Felix Chalumeau, Arthur Flajolet, Geoffrey Cideron, Karim Beguir, Antoine Cully, Olivier Sigaud, and Nicolas Perrin-Gilbert · 2022
Earlier work this paper cites.
Approximating gradients for differentiable quality diversity in reinforcement learning
Bryon Tjanaka, Matthew C Fontaine, Julian Togelius, and Stefanos Nikolaidis · 2022
Earlier work this paper cites.
Continuously discovering novel strategies via reward-switching policy optimization
Zihan Zhou, Wei Fu, Bingliang Zhang, and Yi Wu · 2022
Earlier work this paper cites.
Proximal policy gradient arborescence for quality diversity reinforcement learning
Sumeet Batra, Bryon Tjanaka, Matthew C Fontaine, Aleksei Petrenko, Stefanos Nikolaidis, and Gaurav Sukhatme · 2023
Earlier work this paper cites.
Quality diversity through human feedback: Towards open-ended diversity-driven optimization
Li Ding, Jenny Zhang, Jeff Clune, Lee Spector, and Joel Lehman · 2023
Earlier work this paper cites.
Understanding the effects of rlhf on llm generalisation and diversity
Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica · 2023
Earlier work this paper cites.
Let’s verify step by step
Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2023
Cited alongside, same era.
Deepscaler: Holistic autoscaling for microservices based on spatiotemporal gnn with adaptive graph learning
Chunyang Meng, Shijie Song, Haogang Tong, Maolin Pan, and Yang Yu · 2023
Cited alongside, same era.
Math-shepherd: Verify and reinforce llms step-by-step without human annotations
Peiyi Wang, Lei Li, Zhihong Shao, RX Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, and Zhifang Sui · 2023
Cited alongside, same era.
Quality-similar diversity via population based reinforcement learning
Shuang Wu, Jian Yao, Haobo Fu, Ye Tian, Chao Qian, Yaodong Yang, Qiang Fu, and Yang Wei · 2023
Cited alongside, same era.
Policy space diversity for non-transitive games
Jian Yao, Weiming Liu, Haobo Fu, Yaodong Yang, Stephen McAleer, Qiang Fu, and Wei Yang · 2023
Cited alongside, same era.
Zhiwei He, Tian Liang, Jiahao Xu, Qiuzhi Liu, Xingyu Chen, Yue Wang, Linfeng Song, Dian Yu, Zhenwen Liang, Wenxuan Wang, et al · 2025
Closest in time.
A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility
Andreas Hochlehnert, Hardik Bhatnagar, Vishaal Udandarao, Samuel Albanie, Ameya Prabhu, and Matthias Bethge · 2025
Closest in time.
Curatedthoughts: Data curation for rl training datasets, 2025
Andreas Hochlehnert, Hardik Bhatnagar, Vishaal Udandarao, Ameya Prabhu, and Matthias Bethge · 2025
Closest in time.
Reinforce++: A simple and efficient approach for aligning large language models
Jian Hu · 2025
Closest in time.
Open-reasoner-zero: An open source approach to scaling up reinforcement learning on the base model, 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Leng Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, et al · 2024
Cited alongside, same era.
Vineppo: Unlocking rl potential for llm reasoning through refined credit assignment
Amirhossein Kazemnejad, Milad Aghajohari, Eva Portelance, Alessandro Sordoni, Siva Reddy, Aaron Courville, and Nicolas Le Roux · 2024
Cited alongside, same era.
One fish, two fish, but not the whole sea: Alignment reduces language models’ conceptual diversity
Sonia K Murthy, Tomer Ullman, and Jennifer Hu · 2024
Cited alongside, same era.
Learning to reason with llms
OpenAI · 2024
Cited alongside, same era.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Y Wu, et al · 2024
Cited alongside, same era.
Mathscale: Scaling instruction tuning for mathematical reasoning
Zhengyang Tang, Xingxing Zhang, Benyou Wang, and Furu Wei · 2024
Cited alongside, same era.
Llm-topla: Efficient llm ensemble by maximising diversity
Selim Furkan Tekin, Fatih Ilhan, Tiansheng Huang, Sihao Hu, and Ling Liu · 2024
Cited alongside, same era.
Jingcheng Hu, Yinmin Zhang, Qi Han, Daxin Jiang, Xiangyu Zhang, and Heung-Yeung Shum · 2025
Closest in time.
Evaluation of large language models as solution generators in complex optimization
Beichen Huang, Xingyu Wu, Yu Zhou, Jibin Wu, Liang Feng, Ran Cheng, and Kay Chen Tan · 2025
Closest in time.
How multimodal integration boost the performance of llm for optimization: Case study on capacitated vehicle routing problems
Yuxiao Huang, Wenjie Zhang, Liang Feng, Xingyu Wu, and Kay Chen Tan · 2025
Closest in time.
Open r1: A fully open reproduction of deepseek-r1, January 2025
Hugging Face · 2025
Closest in time.
Limr: Less is more for rl scaling
Xuefeng Li, Haoyang Zou, and Pengfei Liu · 2025
Closest in time.
Preserving diversity in supervised fine-tuning of large language models
Ziniu Li, Congliang Chen, Tian Xu, Zeyu Qin, Jiancong Xiao, Zhi-Quan Luo, and Ruoyu Sun · 2025
Closest in time.
Understanding r1-zero-like training: A critical perspective
Zichen Liu, Changyu Chen, Wenjun Li, Penghui Qi, Tianyu Pang, Chao Du, Wee Sun Lee, and Min Lin · 2025
Closest in time.
Kimi k1.5: Scaling reinforcement learning with llms
Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, et al · 2025
Closest in time.
Diverse policies recovering via pointwise mutual information weighted imitation learning
Hanlin Yang, Jian Yao, Weiming Liu, Qing Wang, Hanmin Qin, Kirk Tang, Jiechao Xiong, Chao Yu, Kai Li, Junliang Xing, et al · 2025
Closest in time.
Dapo: An open-source llm reinforcement learning system at scale, 2025
Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Tiantian Fan, Gaohong Liu, Lingjun Liu, Xin Liu, et al · 2025
Closest in time.
Vapo: Efficient and reliable reinforcement learning for advanced reasoning tasks
Yufeng Yuan, Qiying Yu, Xiaochen Zuo, Ruofei Zhu, Wenyuan Xu, Jiaze Chen, Chengyi Wang, TianTian Fan, Zhengyin Du, Xiangpeng Wei, et al · 2025
Closest in time.
What’s behind ppo’s collapse in long-cot? value optimization holds the secret
Yufeng Yuan, Yu Yue, Ruofei Zhu, Tiantian Fan, and Lin Yan · 2025
Closest in time.
Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?
Yang Yue, Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang, Shiji Song, and Gao Huang · 2025
Closest in time.
Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild, 2025
Weihao Zeng, Yuzhen Huang, Qian Liu, Wei Liu, Keqing He, Zejun Ma, and Junxian He · 2025
Closest in time.
7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient
Weihao Zeng, Yuzhen Huang, Wei Liu, Keqing He, Qian Liu, Zejun Ma, and Junxian He · 2025
Closest in time.
Chong Zhang, Yue Deng, Xiang Lin, Bin Wang, Dianwen Ng, Hai Ye, Xingxuan Li, Yao Xiao, Zhanfeng Mo, Qi Zhang, et al · 2025
Closest in time.
Srpo: A cross-domain implementation of large-scale reinforcement learning on llm
Xiaojiang Zhang, Jinghui Wang, Zifei Cheng, Wenhao Zhuang, Zheng Lin, Minglei Zhang, Shaojie Wang, Yinghan Cui, Chao Wang, Junyi Peng, et al · 2025
Closest in time.
Ttrl: Test-time reinforcement learning
Yuxin Zuo, Kaiyan Zhang, Shang Qu, Li Sheng, Xuekai Zhu, Biqing Qi, Youbang Sun, Ganqu Cui, Ning Ding, and Bowen Zhou · 2025
Closest in time.