Fetching the paper…
Reading the bibliography…
Enhancing the reasoning capabilities of large language models (LLMs) typically relies on massive computational resources and extensive datasets, limiting accessibility for resource-constrained settings.
Explain Yourself! Leveraging Language Models for Commonsense Reasoning
Rajani, N. F.; McCann, B.; Xiong, C.; and Socher, R. 2019 · 2019
Earlier work this paper cites.
Training Verifiers to Solve Math Word Problems
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; Hesse, C.; and Schulman, J. 2021 · 2021
Earlier work this paper cites.
Measuring Mathematical Problem Solving With the MATH Dataset
Hendrycks, D.; Burns, C.; Kadavath, S.; Arora, A.; Basart, S.; Tang, E.; Song, D.; and Steinhardt, J. 2021 · 2021
Earlier work this paper cites.
Show Your Work: Scratchpads for Intermediate Computation with Language Models
Nye, M.; Andreassen, A. J.; Gur-Ari, G.; Michalewski, H.; Austin, J.; Bieber, D.; Dohan, D.; Lewkowycz, A.; Bosma, M.; Luan, D.; Sutton, C.; and Odena, A. 2021 · 2021
Earlier work this paper cites.
Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm
Reynolds, L.; and McDonell, K. 2021 · 2021
Earlier work this paper cites.
Large Language Models are Zero-Shot Reasoners
Kojima, T.; Gu, S. S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y. 2022 · 2022
Earlier work this paper cites.
Solving math word problems with process-and outcome-based feedback
Uesato, J.; Kushman, N.; Kumar, R.; Song, F.; Siegel, N.; Wang, L.; Creswell, A.; Irving, G.; and Higgins, I. 2022 · 2022
Earlier work this paper cites.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; ichter, b.; Xia, F.; Chi, E.; Le, Q. V.; and Zhou, D. 2022 · 2022
Earlier work this paper cites.
STaR: Bootstrapping Reasoning With Reasoning
Zelikman, E.; Wu, Y.; Mu, J.; and Goodman, N. 2022 · 2022
Earlier work this paper cites.
LightEval: A lightweight framework for LLM evaluation
Fourrier, C.; Habib, N.; Kydlíček, H.; Wolf, T.; and Tunstall, L. 2023 · 2023
Earlier work this paper cites.
Self-Refine: Iterative Refinement with Self-Feedback
Madaan, A.; Tandon, N.; Gupta, P.; Hallinan, S.; Gao, L.; Wiegreffe, S.; Alon, U.; Dziri, N.; Prabhumoye, S.; Yang, Y.; Gupta, S.; Majumder, B. P.; Hermann, K.; Welleck, S.; Yazdanbakhsh, A.; and Clark, P. 2023 · 2023
Earlier work this paper cites.
Reflexion: Language Agents with Verbal Reinforcement Learning
Shinn, N.; Cassano, F.; Berman, E.; Gopinath, A.; Narasimhan, K.; and Yao, S. 2023 · 2023
Earlier work this paper cites.
Math-Shepherd: A Label-Free Step-by-Step Verifier for LLMs in Mathematical Reasoning
Wang, P.; Li, L.; Shao, Z.; Xu, R.; Dai, D.; Li, Y.; Chen, D.; Wu, Y.; and Sui, Z. 2023 · 2023
Earlier work this paper cites.
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Zhong, W.; Cui, R.; Guo, Y.; Liang, Y.; Lu, S.; Wang, Y.; Saied, A.; Chen, W.; and Duan, N. 2023 · 2023
Cited alongside, same era.
Introducing Llama 3.1: Our most capable models to date
AI, M. 2024a · 2024
Cited alongside, same era.
Introducing OpenAI o1-preview
AI, O. 2024b · 2024
Cited alongside, same era.
Claude 3.5 Sonnet
Anthropic. 2024 · 2024
Cited alongside, same era.
Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training
Feng, X.; Wan, Z.; Wen, M.; McAleer, S. M.; Wen, Y.; Zhang, W.; and Wang, J. 2024 · 2024
Cited alongside, same era.
Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
Xin, H.; Ren, Z. Z.; Song, J.; Shao, Z.; Zhao, W.; Wang, H.; Liu, B.; Zhang, L.; Lu, X.; Du, Q.; Gao, W.; Zhu, Q.; Yang, D.; Gou, Z.; Wu, Z. F.; Luo, F.; and Ruan, C. 2024 · 2024
Later among the works it cites.
Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Yang, A.; Zhang, B.; Hui, B.; Gao, B.; Yu, B.; Li, C.; Liu, D.; Tu, J.; Zhou, J.; Lin, J.; Lu, K.; Xue, M.; Lin, R.; Liu, T.; Ren, X.; and Zhang, Z. 2024 · 2024
Later among the works it cites.
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Chu, T.; Zhai, Y.; Yang, J.; Tong, S.; Xie, S.; Schuurmans, D.; Le, Q. V.; Levine, S.; and Ma, Y. 2025 · 2025
Closest in time.
Process Reinforcement through Implicit Rewards
Cui, G.; Yuan, L.; Wang, Z.; Wang, H.; Li, W.; He, B.; Fan, Y.; Yu, T.; Xu, Q.; Chen, W.; Yuan, J.; Chen, H.; Zhang, K.; Lv, X.; Wang, S.; Yao, Y.; Han, X.; Peng, H.; Cheng, Y.; Liu, Z.; Sun, M.; Zhou, B.; and Ding, N. 2025 · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gao, B.; Song, F.; Yang, Z.; Cai, Z.; Miao, Y.; Dong, Q.; Li, L.; Ma, C.; Chen, L.; Xu, R.; Tang, Z.; Wang, B.; Zan, D.; Quan, S.; Zhang, G.; Sha, L.; Zhang, Y.; Ren, X.; Liu, T.; and Chang, B. 2024 · 2024
Cited alongside, same era.
Our next-generation model: Gemini 1.5
Google. 2024 · 2024
Cited alongside, same era.
OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
He, C.; Luo, R.; Bai, Y.; Hu, S.; Thai, Z.; Shen, J.; Hu, J.; Han, X.; Huang, Y.; Zhang, Y.; Liu, J.; Qi, L.; Liu, Z.; and Sun, M. 2024 · 2024
Cited alongside, same era.
OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI
Huang, Z.; Wang, Z.; Xia, S.; Li, X.; Zou, H.; Xu, R.; Fan, R.-Z.; Ye, L.; Chern, E.; Ye, Y.; Zhang, Y.; Yang, Y.; Wu, T.; Wang, B.; Sun, S.; Xiao, Y.; Li, Y.; Zhou, F.; Chern, S.; Qin, Y.; Ma, Y.; Su, J.; Liu, Y.; Zheng, Y.; Zhang, S.; Lin, D.; Qiao, Y.; and Liu, P. 2024 · 2024
Cited alongside, same era.
Training language models to self-correct via reinforcement learning
Kumar, A.; Zhuang, V.; Agarwal, R.; Su, Y.; Co-Reyes, J. D.; Singh, A.; Baumli, K.; Iqbal, S.; Bishop, C.; Roelofs, R.; et al. 2024 · 2024
Cited alongside, same era.
NuminaMath
LI, J.; Beeching, E.; Tunstall, L.; Lipkin, B.; Soletskyi, R.; Huang, S. C.; Rasul, K.; Yu, L.; Jiang, A.; Shen, Z.; Qin, Z.; Dong, B.; Zhou, L.; Fleureau, Y.; Lample, G.; and Polu, S. 2024 · 2024
Cited alongside, same era.
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Shao, Z.; Wang, P.; Zhu, Q.; Xu, R.; Song, J.; Bi, X.; Zhang, H.; Zhang, M.; Li, Y. K.; Wu, Y.; and Guo, D. 2024 · 2024
Cited alongside, same era.
Closest in time.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-AI. 2025 · 2025
Closest in time.
Open R1: A fully open reproduction of DeepSeek-R1
Face, H. 2025 · 2025
Closest in time.
rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
Guan, X.; Zhang, L. L.; Liu, Y.; Shang, N.; Sun, Y.; Zhu, Y.; Yang, F.; and Yang, M. 2025 · 2025
Closest in time.
DeepScaleR: Surpassing O1-Preview with a 1.5B Model by Scaling RL
Luo, M.; Tan, S.; Wong, J.; Shi, X.; Tang, W. Y.; Roongta, M.; Cai, C.; Luo, J.; Zhang, T.; Li, L. E.; Popa, R. A.; and Stoica, I. 2025 · 2025
Closest in time.
Muennighoff, N.; Yang, Z.; Shi, W.; Li, X. L.; Fei-Fei, L.; Hajishirzi, H.; Zettlemoyer, L.; Liang, P.; Candès, E.; and Hashimoto, T. 2025 · 2025
Closest in time.
LIMO: Less is More for Reasoning
Ye, Y.; Huang, Z.; Xiao, Y.; Chern, E.; Xia, S.; and Liu, P. 2025 · 2025
Closest in time.
Demystifying Long Chain-of-Thought Reasoning in LLMs
Yeo, E.; Tong, Y.; Niu, X.; Neubig, G.; and Yue, X. 2025 · 2025
Closest in time.
7B Model and 8K Examples: Emerging Reasoning with Reinforcement Learning is Both Effective and Efficient
Zeng, W.; Huang, Y.; Liu, W.; He, K.; Liu, Q.; Ma, Z.; and He, J. 2025 · 2025
Closest in time.