Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have demonstrated their remarkable capacity across a variety of tasks.
Bandit based monte-carlo planning
Kocsis, L.; and Szepesvári, C. 2006 · 2006
Earlier work this paper cites.
A survey of Monte Carlo tree search methods
Browne, C. B.; Powley, E.; Whitehouse, D.; Lucas, S. M.; Cowling, P. I.; Rohlfshagen, P.; Tavener, S.; Perez, D.; Samothrakis, S.; and Colton, S. 2012 · 2012
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; et al. 2021 · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D.; Burns, C.; Kadavath, S.; Arora, A.; Basart, S.; Tang, E.; Song, D.; and Steinhardt, J. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
Nye, M.; Andreassen, A. J.; Gur-Ari, G.; Michalewski, H.; Austin, J.; Bieber, D.; Dohan, D.; Lewkowycz, A.; Bosma, M.; Luan, D.; et al. 2021 · 2021
Earlier work this paper cites.
Complexity-based prompting for multi-step reasoning
Fu, Y.; Peng, H.; Sabharwal, A.; Clark, P.; and Khot, T. 2022 · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Kojima, T.; Gu, S. S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022 · 2022
Earlier work this paper cites.
Large language models still can’t plan (a benchmark for LLMs on planning and reasoning about change)
Valmeekam, K.; Olmo, A.; Sreedharan, S.; and Kambhampati, S. 2022 · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Wang, X.; Wei, J.; Schuurmans, D.; Le, Q.; Chi, E.; Narang, S.; Chowdhery, A.; and Zhou, D. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022 · 2022
Cited alongside, same era.
Star: Bootstrapping reasoning with reasoning
Zelikman, E.; Wu, Y.; Mu, J.; and Goodman, N. 2022 · 2022
Cited alongside, same era.
Reasoning with language model is planning with world model
Hao, S.; Gu, Y.; Ma, H.; Hong, J. J.; Wang, Z.; Wang, D. Z.; and Hu, Z. 2023 · 2023
Cited alongside, same era.
Towards Reasoning in Large Language Models: A Survey
Huang, J.; and Chang, K. C.-C. 2023 · 2023
Cited alongside, same era.
Lightman, H.; Kosaraju, V.; Burda, Y.; Edwards, H.; Baker, B.; Lee, T.; Leike, J.; Schulman, J.; Sutskever, I.; and Cobbe, K. 2023 · 2023
Cited alongside, same era.
Common 7b language models already possess strong math capabilities
Li, C.; Wang, W.; Hu, J.; Wei, Y.; Zheng, N.; Hu, H.; Zhang, Z.; and Peng, H. 2024 · 2024
Later among the works it cites.
Augmenting math word problems via iterative question composing
Liu, H.; Zhang, Y.; Luo, Y.; and Yao, A. C.-C. 2024 · 2024
Later among the works it cites.
Lu, Z.; Zhou, A.; Ren, H.; Wang, K.; Shi, W.; Pan, J.; Zhan, M.; and Li, H. 2024 · 2024
Later among the works it cites.
Improve Mathematical Reasoning in Language Models by Automated Process Supervision
Luo, L.; Liu, Y.; Liu, R.; Phatale, S.; Lara, H.; Li, Y.; Shu, L.; Zhu, Y.; Meng, L.; Sun, J.; et al. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yu, L.; Jiang, W.; Shi, H.; Yu, J.; Liu, Z.; Zhang, Y.; Kwok, J. T.; Li, Z.; Weller, A.; and Liu, W. 2023 · 2023
Cited alongside, same era.
Scaling relationship on learning mathematical reasoning with large language models
Yuan, Z.; Yuan, H.; Li, C.; Dong, G.; Lu, K.; Tan, C.; Zhou, C.; and Zhou, J. 2023 · 2023
Cited alongside, same era.
A survey of large language models
Zhao, W. X.; Zhou, K.; Li, J.; Tang, T.; Wang, X.; Hou, Y.; Min, Y.; Zhang, B.; Zhang, J.; Dong, Z.; et al. 2023 · 2023
Cited alongside, same era.
AlphaMath Almost Zero: process Supervision without process
Chen, G.; Liao, M.; Li, C.; and Fan, K. 2024 · 2024
Cited alongside, same era.
Self-Explore: Enhancing Mathematical Reasoning in Language Models with Fine-grained Rewards
Hwang, H.; Kim, D.; Kim, S.; Ye, S.; and Seo, M. 2024 · 2024
Cited alongside, same era.
Step-dpo: Step-wise preference optimization for long-chain reasoning of llms
Lai, X.; Tian, Z.; Chen, Y.; Yang, S.; Peng, X.; and Jia, J. 2024 · 2024
Cited alongside, same era.
Math-shepherd: Verify and reinforce llms step-by-step without human annotations
Wang, P.; Li, L.; Shao, Z.; Xu, R.; Dai, D.; Li, Y.; Chen, D.; Wu, Y.; and Sui, Z. 2024a
Cited in the paper.
Pang, R. Y.; Yuan, W.; Cho, K.; He, H.; Sukhbaatar, S.; and Weston, J. 2024 · 2024
Later among the works it cites.
Mutual reasoning makes smaller llms stronger problem-solvers
Qi, Z.; Ma, M.; Xu, J.; Zhang, L. L.; Yang, F.; and Yang, M. 2024 · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R.; Sharma, A.; Mitchell, E.; Manning, C. D.; Ermon, S.; and Finn, C. 2024 · 2024
Later among the works it cites.
Reft: Reasoning with reinforced fine-tuning
Trung, L.; Zhang, X.; Jie, Z.; Sun, P.; Jin, X.; and Li, H. 2024 · 2024
Later among the works it cites.
Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Xie, Y.; Goyal, A.; Zheng, W.; Kan, M.-Y.; Lillicrap, T. P.; Kawaguchi, K.; and Shieh, M. 2024 · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T.; Cao, Y.; and Narasimhan, K. 2024 · 2024
Later among the works it cites.