Fetching the paper…
Reading the bibliography…
A promising approach for improving reasoning in large language models is to use process reward models (PRMs).
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
A. Y. Ng, D. Harada, and S. Russell · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour · 1999
Earlier work this paper cites.
Discriminative reranking for natural language parsing
M. Collins · 2000
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
S. Kakade and J. Langford · 2002
Earlier work this paper cites.
Reinforcement and imitation learning via interactive no-regret learning
S. Ross and J. A. Bagnell · 2014
Earlier work this paper cites.
Learning to search better than your teacher
K.-W. Chang, A. Krishnamurthy, A. Agarwal, H. Daumé III, and J. Langford · 2015
Earlier work this paper cites.
Made: Masked autoencoder for distribution estimation
M. Germain, K. Gregor, I. Murray, and H. Larochelle · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton · 2015
Earlier work this paper cites.
A. A. Rusu, S. G. Colmenarejo, C. Gulcehre, G. Desjardins, J. Kirkpatrick, R. Pascanu, V. Mnih, K. Kavukcuoglu, and R. Hadsell · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Earlier work this paper cites.
Thinking fast and slow with deep learning and tree search
T. Anthony, Z. Tian, and D. Barber · 2017
Earlier work this paper cites.
Deeply aggrevated: Differentiable imitation learning for sequential prediction
W. Sun, A. Venkatraman, G. J. Gordon, B. Boots, and J. A. Bagnell · 2017
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
A. Agarwal, S. M. Kakade, J. D. Lee, and G. Mahajan · 2021
Cited alongside, same era.
Measuring mathematical problem solving with the math dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Cited alongside, same era.
Webgpt: Browser-assisted question-answering with human feedback
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, et al · 2021
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
Solving math word problems with process-and outcome-based feedback
J. Uesato, N. Kushman, R. Kumar, F. Song, N. Siegel, L. Wang, A. Creswell, G. Irving, and I. Higgins · 2022
Gemma: Open models based on gemini research and technology
Gemma Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivière, M. S. Kale, J. Love, et al · 2024
Closest in time.
Teaching large language models to reason with reinforcement learning
A. Havrilla, Y. Du, S. C. Raparthy, C. Nalmpantis, J. Dwivedi-Yu, M. Zhuravinskyi, E. Hambro, S. Sukhbaatar, and R. Raileanu · 2024
Closest in time.
V-star: Training verifiers for self-taught reasoners
A. Hosseini, X. Yuan, N. Malkin, A. Courville, A. Sordoni, and R. Agarwal · 2024
Closest in time.
H. Hwang, D. Kim, S. Kim, S. Ye, and M. Seo · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Star: Bootstrapping reasoning with reasoning
E. Zelikman, Y. Wu, J. Mu, and N. Goodman · 2022
Cited alongside, same era.
Learning to generate better than your llm
J. D. Chang, K. Brantley, R. Ramamurthy, D. Misra, and W. Sun · 2023
Cited alongside, same era.
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe · 2023
Cited alongside, same era.
Let’s reward step by step: Step-level reward model as the navigators for reasoning
Q. Ma, H. Zhou, T. Liu, J. Yuan, P. Liu, Y. You, and H. Yang · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn · 2023
Cited alongside, same era.
Progprompt: Generating situated robot task plans using large language models
I. Singh, V. Blukis, A. Mousavian, A. Goyal, D. Xu, J. Tremblay, D. Fox, J. Thomason, and A. Garg · 2023
Cited alongside, same era.
Outcome-supervised verifiers for planning in mathematical reasoning
F. Yu, A. Gao, and B. Wang · 2023
Cited alongside, same era.
A. Kazemnejad, M. Aghajohari, E. Portelance, A. Sordoni, S. Reddy, A. Courville, and N. L. Roux · 2024
Closest in time.
Improve mathematical reasoning in language models by automated process supervision
L. Luo, Y. Liu, R. Liu, S. Phatale, H. Lara, Y. Li, L. Shu, Y. Zhu, L. Meng, J. Sun, et al · 2024
Closest in time.
Rl on incorrect synthetic data scales the efficiency of llm math reasoning by eight-fold
A. Setlur, S. Garg, X. Geng, N. Garg, V. Smith, and A. Kumar · 2024
Closest in time.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, M. Zhang, Y. Li, Y. Wu, and D. Guo · 2024
Closest in time.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
C. Snell, J. Lee, K. Xu, and A. Kumar · 2024
Closest in time.
Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data, 2024
F. Tajwar, A. Singh, A. Sharma, R. Rafailov, J. Schneider, T. Xie, S. Ermon, C. Finn, and A. Kumar · 2024
Closest in time.
Math-shepherd: Verify and reinforce llms step-by-step without human annotations, 2024
P. Wang, L. Li, Z. Shao, R. X. Xu, D. Dai, Y. Li, D. Chen, Y. Wu, and Z. Sui · 2024
Closest in time.
An empirical analysis of compute-optimal inference for problem-solving with language models
Y. Wu, Z. Sun, S. Li, S. Welleck, and Y. Yang · 2024
Closest in time.
Generative verifiers: Reward modeling as next-token prediction
L. Zhang, A. Hosseini, H. Bansal, M. Kazemi, A. Kumar, and R. Agarwal · 2024
Closest in time.