Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are increasingly applied to complex reasoning tasks that require executing several complex steps before receiving any reward.
Fine-tuning Language Models from Human Preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P. F., and Irving, G · 1909
Earlier work this paper cites.
Introduction to Reinforcement Learning
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Policy Gradient Methods for Reinforcement Learning with Function Approximation
Sutton, R. S., McAllester, D. A., Singh, S., and Mansour, Y · 1999
Earlier work this paper cites.
Variance Reduction Techniques for Gradient Estimates in Reinforcement Learning
Greensmith, E., Bartlett, P. L., and Baxter, J · 2001
Earlier work this paper cites.
Trust Region Policy Optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M. I., and Moritz, P · 2015
Earlier work this paper cites.
High-dimensional Continuous Control Using Generalized Advantage Estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M. I., and Abbeel, P · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T. P., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D · 2016
Earlier work this paper cites.
Thinking Fast and Slow with Deep Learning and Tree Search
Anthony, T., Tian, Z., and Barber, D · 2017
Earlier work this paper cites.
Proximal Policy Optimization Algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Mastering Chess and Shogi by Self-play with a General Reinforcement Learning Algorithm
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T. P., Simonyan, K., and Hassabis, D · 2017
Earlier work this paper cites.
Training Verifiers to Solve Math Word Problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Measuring Mathematical Problem Solving With the MATH Dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Solving Quantitative Reasoning Problems with Language Models
Lewkowycz, A., Andreassen, A., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V. V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., Wu, Y., Neyshabur, B., Gur-Ari, G., and Misra, V · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., and Lowe, R · 2022
Earlier work this paper cites.
Solving math word problems with process- and outcome-based feedback
Uesato, J., Kushman, N., Kumar, R., Song, H. F., Siegel, N. Y., Wang, L., Creswell, A., Irving, G., and Higgins, I · 2022
Earlier work this paper cites.
Learning to generate better than your llm
Chang, J. D., Brantley, K., Ramamurthy, R., Misra, D., and Sun, W · 2023
Earlier work this paper cites.
Reasoning with Language Model is Planning with World Model
Hao, S., Gu, Y., Ma, H., Hong, J. J., Wang, Z., Wang, D. Z., and Hu, Z · 2023
Earlier work this paper cites.
Efficient Memory Management for Large Language Model Serving with PagedAttention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Cited alongside, same era.
Let’s reward step by step: Step-level reward model as the Navigators for Reasoning
Ma, Q., Zhou, H., Liu, T., Yuan, J., Liu, P., You, Y., and Yang, H · 2023
Cited alongside, same era.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2023
Cited alongside, same era.
Llama 2: Open Foundation and Fine-tuned Chat Models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Canton-Ferrer, C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., et al · 2023
Cited alongside, same era.
Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
Agent q: Advanced reasoning and learning for autonomous ai agents, 2024
Putta, P., Mills, E., Garg, N., Motwani, S., Finn, C., Garg, D., and Rafailov, R · 2024
Closest in time.
Qwen2.5-Math: The world’s leading open-sourced mathematical LLMs
Qwen · 2024
Closest in time.
Notes on the KL-divergence Approximation
Schulman, J · 2024
Closest in time.
RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-fold
Setlur, A., Garg, S., Geng, X., Garg, N., Smith, V., and Kumar, A · 2024
Closest in time.
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Zhang, M., Li, Y. K., Wu, Y., and Guo, D · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuan, Z., Yuan, H., Li, C., Dong, G., Tan, C., and Zhou, C · 2023
Cited alongside, same era.
Secrets of RLHF in Large Language Models Part I: PPO
Zheng, R., Dou, S., Gao, S., Hua, Y., Shen, W., Wang, B., Liu, Y., Jin, S., Liu, Q., Zhou, Y., Xiong, L., Chen, L., Xi, Z., Xu, N., Lai, W., Zhu, M., Chang, C., Yin, Z., Weng, R., Cheng, W., Huang, H., Sun, T., Yan, H., Gui, T., Zhang, Q., Qiu, X., and Huang, X · 2023
Cited alongside, same era.
Back to Basics: Revisiting REINFORCE-style Optimization for Learning from Human Feedback in LLMs
Ahmadian, A., Cremer, C., Gallé, M., Fadaee, M., Kreutzer, J., Pietquin, O., Üstün, A., and Hooker, S · 2024
Cited alongside, same era.
LoRA Learns Less and Forgets Less
Biderman, D., Ortiz, J. J. G., Portes, J., Paul, M., Greengard, P., Jennings, C., King, D., Havens, S., Chiley, V., Frankle, J., Blakeney, C., and Cunningham, J. P · 2024
Cited alongside, same era.
AlphaMath Almost Zero: process Supervision without process
Chen, G., Liao, M., Li, C., and Fan, K · 2024
Cited alongside, same era.
The N+ Implementation Details of RLHF with PPO: A Case Study on TL;DR Summarization
Huang, S., Noukhovitch, M., Hosseini, A., Rasul, K., Wang, W., and Tunstall, L · 2024
Cited alongside, same era.
Hwang, H., Kim, D., Kim, S., Ye, S., and Seo, M · 2024
Cited alongside, same era.
Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback
Ivison, H., Wang, Y., Liu, J., Wu, Z., Pyatkin, V., Lambert, N., Smith, N. A., Choi, Y., and Hajishirzi, H · 2024
Cited alongside, same era.
Beyond Human Data: Scaling Self-training for Problem-solving with Language Models
Singh, A., Co-Reyes, J. D., Agarwal, R., Anand, A., Patil, P., Garcia, X., Liu, P. J., Harrison, J., Lee, J., Xu, K., Parisi, A. T., Kumar, A., Alemi, A. A., Rizkowsky, A., Nova, A., Adlam, B., Bohnet, B., Elsayed, G. F., Sedghi, H., Mordatch, I., Simpson, I., Gur, I., Snoek, J., Pennington, J., Hron, J., Kenealy, K., Swersky, K., Mahajan, K., Culp, L., Xiao, L., Bileschi, M. L., Constant, N., Novak, R., Liu, R., Warkentin, T., Qian, Y., Bansal, Y., Dyer, E., Neyshabur, B., Sohl-Dickstein, J., and Fiedel, N · 2024
Closest in time.
Gymnasium: A standard interface for reinforcement learning environments
Towers, M., Kwiatkowski, A., Terry, J., Balis, J. U., De Cola, G., Deleu, T., Goulão, M., Kallinteris, A., Krimmel, M., KG, A., et al · 2024
Closest in time.
ReFT: Reasoning with Reinforced Fine-tuning
Trung, L. Q., Zhang, X., Jie, Z., Sun, P., Jin, X., and Li, H · 2024
Closest in time.
AlphaZero-like Tree-search can Guide Large Language Model Decoding and Training
Wan, Z., Feng, X., Wen, M., McAleer, S. M., Wen, Y., Zhang, W., and Wang, J · 2024
Closest in time.
Math-shepherd: Verify and reinforce llms step-by-step without human annotations
Wang, P., Li, L., Shao, Z., Xu, R. X., Dai, D., Li, Y., Chen, D., Wu, Y., and Sui, Z · 2024
Closest in time.
Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Xie, Y., Goyal, A., Zheng, W., Kan, M., Lillicrap, T. P., Kawaguchi, K., and Shieh, M · 2024
Closest in time.
Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
Xu, S., Fu, W., Gao, J., Ye, W., Liu, W., Mei, Z., Wang, G., Yu, C., and Wu, Y · 2024
Closest in time.
ReST-MCTS*: LLM Self-training via Process Reward Guided Tree Search
Zhang, D., Zhoubian, S., Yue, Y., Dong, Y., and Tang, J · 2024
Closest in time.
Sglang: Efficient execution of structured language model programs
Zheng, L., Yin, L., Xie, Z., Sun, C., Huang, J., Yu, C. H., Cao, S., Kozyrakis, C., Stoica, I., Gonzalez, J. E., Barrett, C., and Sheng, Y · 2024
Closest in time.
Sft memorizes, rl generalizes: A comparative study of foundation model post-training, 2025
Chu, T., Zhai, Y., Yang, J., Tong, S., Xie, S., Schuurmans, D., Le, Q. V., Levine, S., and Ma, Y · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI, Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., et al · 2025
Closest in time.