Fetching the paper…
Reading the bibliography…
Step-level reward models (SRMs) can significantly enhance mathematical reasoning performance through process supervision or step-level preference alignment based on reinforcement learning.
Training verifiers to solve math word problems
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; et al. 2021 · 2021
Earlier work this paper cites.
Measuring Mathematical Problem Solving With the MATH Dataset
Hendrycks, D.; Burns, C.; Kadavath, S.; Arora, A.; Basart, S.; Tang, E.; Song, D.; and Steinhardt, J. 2021 · 2021
Earlier work this paper cites.
Self-Consistency Improves Chain of Thought Reasoning in Language Models
Wang, X.; Wei, J.; Schuurmans, D.; Le, Q. V.; Chi, E. H.; Narang, S.; Chowdhery, A.; and Zhou, D. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022 · 2022
Earlier work this paper cites.
Llemma: An Open Language Model for Mathematics
Azerbayev, Z.; Schoelkopf, H.; Paster, K.; Dos Santos, M.; McAleer, S. M.; Jiang, A. Q.; Deng, J.; Biderman, S.; and Welleck, S. 2023 · 2023
Earlier work this paper cites.
Everything of thoughts: Defying the law of penrose triangle for thought generation
Ding, R.; Zhang, C.; Wang, L.; Xu, Y.; Ma, M.; Zhang, W.; Qin, S.; Rajmohan, S.; Lin, Q.; and Zhang, D. 2023 · 2023
Earlier work this paper cites.
Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training
Feng, X.; Wan, Z.; Wen, M.; Wen, Y.; Zhang, W.; and Wang, J. 2023 · 2023
Earlier work this paper cites.
Reasoning with Language Model is Planning with World Model
Hao, S.; Gu, Y.; Ma, H.; Hong, J.; Wang, Z.; Wang, D.; and Hu, Z. 2023 · 2023
Earlier work this paper cites.
Let’s Verify Step by Step
Lightman, H.; Kosaraju, V.; Burda, Y.; Edwards, H.; Baker, B.; Lee, T.; Leike, J.; Schulman, J.; Sutskever, I.; and Cobbe, K. 2023 · 2023
Earlier work this paper cites.
Let’s reward step by step: Step-Level reward model as the Navigators for Reasoning
Ma, Q.; Zhou, H.; Liu, T.; Yuan, J.; Liu, P.; You, Y.; and Yang, H. 2023 · 2023
Cited alongside, same era.
A survey of large language models
Zhao, W. X.; Zhou, K.; Li, J.; Tang, T.; Wang, X.; Hou, Y.; Min, Y.; Zhang, B.; Zhang, J.; Dong, Z.; et al. 2023 · 2023
Cited alongside, same era.
Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Zhou, D.; Schärli, N.; Hou, L.; Wei, J.; Scales, N.; Wang, X.; Schuurmans, D.; Cui, C.; Bousquet, O.; Le, Q. V.; et al. 2023 · 2023
Cited alongside, same era.
Graph of thoughts: Solving elaborate problems with large language models
Besta, M.; Blach, N.; Kubicek, A.; Gerstenberger, R.; Podstawski, M.; Gianinazzi, L.; Gajda, J.; Lehmann, T.; Niewiadomski, H.; Nyczyk, P.; et al. 2024 · 2024
Cited alongside, same era.
The Llama 3 Herd of Models
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z.; Wang, P.; Zhu, Q.; Xu, R.; Song, J.; Zhang, M.; Li, Y.; Wu, Y.; and Guo, D. 2024 · 2024
Closest in time.
Math-shepherd: Verify and reinforce llms step-by-step without human annotations
Wang, P.; Li, L.; Shao, Z.; Xu, R.; Dai, D.; Li, Y.; Chen, D.; Wu, Y.; and Sui, Z. 2024 · 2024
Closest in time.
Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Xie, Y.; Goyal, A.; Zheng, W.; Kan, M.-Y.; Lillicrap, T. P.; Kawaguchi, K.; and Shieh, M. 2024 · 2024
Closest in time.
Qwen2 Technical Report
Yang, A.; Yang, B.; Hui, B.; Zheng, B.; Yu, B.; Zhou, C.; Li, C.; Li, C.; Liu, D.; Huang, F.; et al. 2024 · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T.; Cao, Y.; and Narasimhan, K. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; et al. 2024 · 2024
Cited alongside, same era.
Language is primarily a tool for communication rather than thought
Fedorenko, E.; Piantadosi, S. T.; and Gibson, E. A. 2024 · 2024
Cited alongside, same era.
LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models
Hao, S.; Gu, Y.; Luo, H.; Liu, T.; Shao, X.; Wang, X.; Xie, S.; Ma, H.; Samavedhi, A.; Gao, Q.; et al. 2024 · 2024
Cited alongside, same era.
Enhancing Length Generalization for Attention Based Knowledge Tracing Models with Linear Biases
Li, X.; Bai, Y.; Guo, T.; Liu, Z.; Huang, Y.; Zhao, X.; Xia, F.; Luo, W.; and Weng, J. 2024 · 2024
Cited alongside, same era.
AlphaMath Almost Zero: process Supervision without process
Chen, G.; Liao, M.; Li, C.; and Fan, K. 2024a
Cited in the paper.
Step-level Value Preference Optimization for Mathematical Reasoning
Chen, G.; Liao, M.; Li, C.; and Fan, K. 2024b
Cited in the paper.
Knowledge tracing as language processing: A large-scale autoregressive paradigm
Zhan, B.; Guo, T.; Li, X.; Hou, M.; Liang, Q.; Gao, B.; Luo, W.; and Liu, Z. 2024 · 2024
Closest in time.
Zhang, D.; Li, J.; Huang, X.; Zhou, D.; Li, Y.; and Ouyang, W. 2024 · 2024
Closest in time.
Automatic Lesson Plan Generation via Large Language Models with Self-critique Prompting
Zheng, Y.; Li, X.; Huang, Y.; Liang, Q.; Guo, T.; Hou, M.; Gao, B.; Tian, M.; Liu, Z.; and Luo, W. 2024 · 2024
Closest in time.