Fetching the paper…
Reading the bibliography…
The leaderboard of Large Language Models (LLMs) in mathematical tasks has been continuously updated.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2019 · 2019
Earlier work this paper cites.
A Study of Automatic Metrics for the Evaluation of Natural Language Explanations
Clinciu, M.-A.; Eshghi, A.; and Hastie, H. 2021 · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; et al. 2021 · 2021
Earlier work this paper cites.
Measuring Mathematical Problem Solving With the MATH Dataset
Hendrycks, D.; Burns, C.; Kadavath, S.; Arora, A.; Basart, S.; Tang, E.; Song, D.; and Steinhardt, J. 2021 · 2021
Earlier work this paper cites.
ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning
Golovneva, O.; Chen, M. P.; Poff, S.; Corredor, M.; Zettlemoyer, L.; Fazel-Zarandi, M.; and Celikyilmaz, A. 2022 · 2022
Earlier work this paper cites.
Solving quantitative reasoning problems with language models
Lewkowycz, A.; Andreassen, A.; Dohan, D.; Dyer, E.; Michalewski, H.; Ramasesh, V.; Slone, A.; Anil, C.; Schlag, I.; Gutman-Solo, T.; et al. 2022 · 2022
Earlier work this paper cites.
Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought
Saparov, A.; and He, H. 2022 · 2022
Earlier work this paper cites.
Solving math word problems with process-and outcome-based feedback
Uesato, J.; Kushman, N.; Kumar, R.; Song, F.; Siegel, N.; Wang, L.; Creswell, A.; Irving, G.; and Higgins, I. 2022 · 2022
Earlier work this paper cites.
Llemma: An Open Language Model for Mathematics
Azerbayev, Z.; Schoelkopf, H.; Paster, K.; Dos Santos, M.; McAleer, S. M.; Jiang, A. Q.; Deng, J.; Biderman, S.; and Welleck, S. 2023 · 2023
Earlier work this paper cites.
Generative AI for Math: Abel
Chern, E.; Zou, H.; Li, X.; Hu, J.; Feng, K.; Li, J.; and Liu, P. 2023 · 2023
Cited alongside, same era.
Safe RLHF: Safe Reinforcement Learning from Human Feedback
Dai, J.; Pan, X.; Sun, R.; Ji, J.; Xu, X.; Liu, M.; Wang, Y.; and Yang, Y. 2023 · 2023
Cited alongside, same era.
Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D. S.; Casas, D. d. l.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; et al. 2023 · 2023
Cited alongside, same era.
Let’s Verify Step by Step
Lightman, H.; Kosaraju, V.; Burda, Y.; Edwards, H.; Baker, B.; Lee, T.; Leike, J.; Schulman, J.; Sutskever, I.; and Cobbe, K. 2023 · 2023
Cited alongside, same era.
A Survey of Deep Learning for Mathematical Reasoning
Lu, P.; Qiu, L.; Yu, W.; Welleck, S.; and Chang, K.-W. 2023 · 2023
Cited alongside, same era.
Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct
Arb: Advanced reasoning benchmark for large language models
Sawada, T.; Paleka, D.; Havrilla, A.; Tadepalli, P.; Vidas, P.; Kranias, A.; Nay, J. J.; Gupta, K.; and Komatsuzaki, A. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023 · 2023
Later among the works it cites.
LLMs cannot find reasoning errors, but can correct them!
Tyen, G.; Mansoor, H.; Chen, P.; Mak, T.; and Cărbune, V. 2023 · 2023
Later among the works it cites.
Math-Shepherd: A Label-Free Step-by-Step Verifier for LLMs in Mathematical Reasoning
Wang, P.; Li, L.; Shao, Z.; Xu, R.; Dai, D.; Li, Y.; Chen, D.; Wu, Y.; and Sui, Z. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Luo, H.; Sun, Q.; Xu, C.; Zhao, P.; Lou, J.; Tao, C.; Geng, X.; Lin, Q.; Chen, S.; and Zhang, D. 2023 · 2023
Cited alongside, same era.
Let’s reward step by step: Step-Level reward model as the Navigators for Reasoning
Ma, Q.; Zhou, H.; Liu, T.; Yuan, J.; Liu, P.; You, Y.; and Yang, H. 2023 · 2023
Cited alongside, same era.
GPT-4 technical report
OpenAI, R. 2023 · 2023
Cited alongside, same era.
ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness
Prasad, A.; Saha, S.; Zhou, X.; and Bansal, M. 2023 · 2023
Cited alongside, same era.
MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
Yu, L.; Jiang, W.; Shi, H.; Jincheng, Y.; Liu, Z.; Zhang, Y.; Kwok, J.; Li, Z.; Weller, A.; and Liu, W. 2023 · 2023
Later among the works it cites.
Interpretable Math Word Problem Solution Generation via Step-by-step Planning
Zhang, M.; Wang, Z.; Yang, Z.; Feng, W.; and Lan, A. 2023 · 2023
Later among the works it cites.
Hao, S.; Gu, Y.; Luo, H.; Liu, T.; Shao, X.; Wang, X.; Xie, S.; Ma, H.; Samavedhi, A.; Gao, Q.; Wang, Z.; and Hu, Z. 2024 · 2024
Closest in time.
Augmenting math word problems via iterative question composing
Liu, H.; and Yao, A. C.-C. 2024 · 2024
Closest in time.
MR-GSM8K: A Meta-Reasoning Revolution in Large Language Model Evaluation
Zeng, Z.; Chen, P.; Liu, S.; Jiang, H.; and Jia, J. 2024 · 2024
Closest in time.