Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) show great promise in complex reasoning, with Reinforcement Learning with Verifiable Rewards (RLVR) being a key enhancement strategy.
Neuronlike adaptive elements that can solve difficult learning control problems
A. G. Barto, R. S. Sutton, and C. W. Anderson · 1983
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. I. Jordan, and P. Abbeel · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
Trl: Transformer reinforcement learning
L. von Werra, Y. Belkada, L. Tunstall, E. Beeching, T. Thrush, N. Lambert, S. Huang, K. Rasul, and Q. Gallouédec · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Solving quantitative reasoning problems with language models
A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V. Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning
E. Zelikman, Y. Wu, J. Mu, and N. Goodman · 2022
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica · 2023
Earlier work this paper cites.
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe · 2023
Earlier work this paper cites.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2023
Earlier work this paper cites.
Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms
A. Ahmadian, C. Cremer, M. Gallé, M. Fadaee, J. Kreutzer, O. Pietquin, A. Üstün, and S. Hooker · 2024
Cited alongside, same era.
On designing effective rl reward at training time for llm reasoning
J. Gao, S. Xu, W. Ye, W. Liu, C. He, W. Fu, Z. Mei, G. Wang, and Y. Wu · 2024
Cited alongside, same era.
C. He, R. Luo, Y. Bai, S. Hu, Z. L. Thai, J. Shen, J. Hu, X. Han, Y. Huang, Y. Zhang, et al · 2024
Cited alongside, same era.
Self-improvement in language models: The sharpening mechanism
A. Huang, A. Block, D. J. Foster, D. Rohatgi, C. Zhang, M. Simchowitz, J. T. Ash, and A. Krishnamurthy · 2024
Cited alongside, same era.
Open-reasoner-zero: An open source approach to scaling up reinforcement learning on the base model
J. Hu, Y. Zhang, Q. Han, D. Jiang, X. Zhang, and H.-Y. Shum · 2025
Closest in time.
Learning to solve and verify: A self-play framework for code and test generation
Z. Lin, S. Shen, J. Shang, J. Weston, and Y. Nie · 2025
Closest in time.
There may not be aha moment in r1-zero-like training — a pilot study
Z. Liu, C. Chen, W. Li, T. Pang, C. Du, and M. Lin · 2025
Closest in time.
Heimdall: test-time scaling on the generative verification
W. Shi and X. Jin · 2025
Closest in time.
Crossing the reward bridge: Expanding rl with verifiable rewards across diverse domains, 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, et al · 2024
Cited alongside, same era.
Tülu 3: Pushing frontiers in open language model post-training
N. Lambert, J. D. Morrison, V. Pyatkin, S. Huang, H. Ivison, F. Brahman, L. J. V. Miranda, A. Liu, N. Dziri, X. Lyu, Y. Gu, S. Malik, V. Graf, J. D. Hwang, J. Yang, R. L. Bras, O. Tafjord, C. Wilhelm, L. Soldaini, N. A. Smith, Y. Wang, P. Dasigi, and H. Hajishirzi · 2024
Cited alongside, same era.
Gpt-4o, 2024
OpenAI · 2024
Cited alongside, same era.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models, 2024
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. K. Li, Y. Wu, and D. Guo · 2024
Cited alongside, same era.
Hybridflow: A flexible and efficient rlhf framework
G. Sheng, C. Zhang, Z. Ye, X. Wu, W. Zhang, R. Zhang, Y. Peng, H. Lin, and C. Wu · 2024
Cited alongside, same era.
Mind the gap: Examining the self-improvement capabilities of large language models
Y. Song, H. Zhang, C. Eisenach, S. Kakade, D. Foster, and U. Ghai · 2024
Cited alongside, same era.
Enhancing llm reasoning via critique models with test-time and training-time supervision
Z. Xi, D. Yang, J. Huang, J. Tang, G. Li, Y. Ding, W. He, B. Hong, S. Do, W. Zhan, et al · 2024
Cited alongside, same era.
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu · 2024
Cited alongside, same era.
Y. Su, D. Yu, L. Song, J. Li, H. Mi, Z. Tu, M. Zhang, and D. Yu · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
K. Team, A. Du, B. Gao, B. Xing, C. Jiang, C. Chen, C. Li, C. Xiao, C. Du, C. Liao, et al · 2025
Closest in time.
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
S. Toshniwal, W. Du, I. Moshkov, B. Kisacanin, A. Ayrapetyan, and I. Gitman · 2025
Closest in time.
P versus np problem — Wikipedia, the free encyclopedia, 2025
Wikipedia contributors · 2025
Closest in time.
Teaching language models to critique via reinforcement learning
Z. Xie, L. Chen, W. Mao, J. Xu, L. Kong, et al · 2025
Closest in time.
Self-rewarding correction for mathematical reasoning
W. Xiong, H. Zhang, C. Ye, L. Chen, N. Jiang, and T. Zhang · 2025
Closest in time.
Demystifying long chain-of-thought reasoning in llms
E. Yeo, Y. Tong, M. Niu, G. Neubig, and X. Yue · 2025
Closest in time.
Y. Yue, Z. Chen, R. Lu, A. Zhao, Z. Wang, Y. Yue, S. Song, and G. Huang · 2025
Closest in time.
W. Zeng, Y. Huang, Q. Liu, W. Liu, K. He, Z. Ma, and J. He · 2025
Closest in time.
Generative verifiers: Reward modeling as next-token prediction, 2025
L. Zhang, A. Hosseini, H. Bansal, M. Kazemi, A. Kumar, and R. Agarwal · 2025
Closest in time.