Fetching the paper…
Reading the bibliography…
Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for complex reasoning tasks with clear correctness signals such as math and coding.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Earlier work this paper cites.
The false promise of imitating proprietary llms
A. Gudibande, E. Wallace, C. Snell, X. Geng, H. Liu, P. Abbeel, S. Levine, and D. Song · 2023
Earlier work this paper cites.
Let’s verify step by step
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe · 2023
Earlier work this paper cites.
A long way to go: Investigating length correlations in rlhf
P. Singhal, T. Goyal, J. Xu, and G. Durrett · 2023
Earlier work this paper cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al · 2023
Earlier work this paper cites.
H. Hashemi, J. Eisner, C. Rosset, B. Van Durme, and C. Kedzie · 2024
Earlier work this paper cites.
A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al · 2024
Earlier work this paper cites.
A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, et al · 2024
Earlier work this paper cites.
Prometheus 2: An open source language model specialized in evaluating other language models
S. Kim, J. Suk, S. Longpre, B. Y. Lin, J. Shin, S. Welleck, G. Neubig, M. Lee, K. Lee, and M. Seo · 2024
Earlier work this paper cites.
T \ \backslash " ulu 3: Pushing frontiers in open language model post-training
N. Lambert, J. Morrison, V. Pyatkin, S. Huang, H. Ivison, F. Brahman, L. J. V. Miranda, A. Liu, N. Dziri, S. Lyu, et al · 2024
Earlier work this paper cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. Li, Y. Wu, et al · 2024
Earlier work this paper cites.
H. Wang, Y. Lin, W. Xiong, R. Yang, S. Diao, S. Qiu, H. Zhao, and T. Zhang · 2024
Earlier work this paper cites.
Improving reward models with synthetic critiques
Z. Ye, F. Greenlee-Scott, M. Bartolo, P. Blunsom, J. A. Campos, and M. Gallé · 2024
Cited alongside, same era.
R3: Robust rubric-agnostic reward models
D. Anugraha, Z. Tang, L. J. V. Miranda, H. Zhao, M. R. Farhansyah, G. Kuwanto, D. Wijaya, and G. I. Winata · 2025
Cited alongside, same era.
Healthbench: Evaluating large language models towards improved human health
R. K. Arora, J. Wei, R. S. Hicks, P. Bowman, J. Quiñonero-Candela, F. Tsimpourlas, M. Sharman, M. Shah, A. Vallone, A. Beutel, et al · 2025
Cited alongside, same era.
Rm-r1: Reward modeling as reasoning
X. Chen, G. Li, Z. Wang, B. Jin, C. Qian, Y. Wang, H. Wang, Y. Zhang, D. Zhang, T. Zhang, et al · 2025
Cited alongside, same era.
D. Lu, X. Tan, R. Xu, T. Yao, C. Qu, W. Chu, Y. Xu, and Y. Qi · 2025
Closest in time.
General-reasoner: Advancing llm reasoning across all domains
X. Ma, Q. Liu, D. Jiang, G. Zhang, Z. Ma, and W. Chen · 2025
Closest in time.
Openai o3-min, 2025
OpenAI o3-mini · 2025
Closest in time.
Rubric is all you need: Enhancing llm-based code evaluation with question-specific rubrics
A. Pathak, R. Gandhi, V. Uttam, Y. Nakka, A. R. Jindal, P. Ghosh, A. Ramamoorthy, S. Verma, A. Mittal, A. Ased, et al · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Cui, L. Yuan, Z. Wang, H. Wang, W. Li, B. He, Y. Fan, T. Yu, Q. Xu, W. Chen, et al · 2025
Cited alongside, same era.
Qa-lign: Aligning llms through constitutionally decomposed qa
J. Dineen, A. RRV, Q. Liu, Z. Xu, X. Ye, M. Shen, Z. Li, S. Lu, C. Baral, M. Chen, et al · 2025
Cited alongside, same era.
Configurable preference tuning with rubric-guided synthetic data
V. Gallego · 2025
Cited alongside, same era.
General reasoning, 2025
General Reasoning · 2025
Cited alongside, same era.
Process reward models that think
M. Khalifa, R. Agarwal, L. Logeswaran, J. Kim, H. Peng, M. Lee, H. Lee, and L. Wang · 2025
Cited alongside, same era.
Toward evaluative thinking: Meta policy optimization with evolving reward models
Z. M. Kim, C. Park, V. Raheja, S. Kim, and D. Kang · 2025
Cited alongside, same era.
Enhancing reasoning through process supervision with monte carlo tree search
S. Li, S. Dong, K. Luan, X. Di, and C. Ding · 2025
Cited alongside, same era.
Huatuogpt-o1, towards medical complex reasoning with llms
J. Chen, Z. Cai, K. Ji, X. Wang, W. Liu, R. Wang, J. Hou, and B. Wang
Cited in the paper.
J. Ruan, I. Nair, S. Cao, A. Liu, S. Munir, M. Pollens-Dempsey, T. Chiang, L. Kates, N. David, S. Chen, et al · 2025
Closest in time.
V. Sirdeshmukh, K. Deshpande, J. Mols, L. Jin, E.-Y. Cardona, D. Lee, J. Kritz, W. Primack, S. Yue, and C. Xing · 2025
Closest in time.
Checklists are better than reward models for aligning language models
V. Viswanathan, Y. Sun, S. Ma, X. Kong, M. Cao, G. Neubig, and T. Wu · 2025
Closest in time.
J1: Incentivizing thinking in llm-as-a-judge via reinforcement learning
C. Whitehouse, T. Wang, P. Yu, X. Li, J. Weston, I. Kulikov, and S. Saha · 2025
Closest in time.
Naturalreasoning: Reasoning in the wild with 2.8 m challenging questions
W. Yuan, J. Yu, S. Jiang, K. Padthe, Y. Li, I. Kulikov, K. Cho, D. Wang, Y. Tian, J. E. Weston, et al · 2025
Closest in time.
Med-rlvr: Emerging medical reasoning from a 3b base model via reinforcement learning
S. Zhang, Q. Liu, G. Qin, T. Naumann, and H. Poon · 2025
Closest in time.
Learning to reason without external rewards
X. Zhao, Z. Kang, A. Feng, S. Levine, and D. Song · 2025
Closest in time.