Fetching the paper…
Reading the bibliography…
Mathematical reasoning in Large Language Models (LLMs) is often evaluated using benchmarks with limited numerical ranges, failing to reflect real-world problem-solving across diverse scales.
Scaling laws for autoregressive generative modeling
Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B Brown, Prafulla Dhariwal, Scott Gray, et al. 2020 · 2010
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021 · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Earlier work this paper cites.
Text and patterns: For effective chain of thought, it takes two to tango
Aman Madaan and Amir Yazdanbakhsh. 2022 · 2022
Earlier work this paper cites.
A causal framework to quantify the robustness of mathematical reasoning with language models
Alessandro Stolfo, Zhijing Jin, Kumar Shridhar, Bernhard Schölkopf, and Mrinmaya Sachan. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought hub: A continuous effort to measure large language models’ reasoning performance
Yao Fu, Litu Ou, Mingyu Chen, Yuhao Wan, Hao Peng, and Tushar Khot. 2023 · 2023
Earlier work this paper cites.
Limitations of language models in arithmetic and symbolic induction
Jing Qian, Hong Wang, Zekun Li, Shiyang Li, and Xifeng Yan. 2023 · 2023
Earlier work this paper cites.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023 · 2023
Earlier work this paper cites.
An independent evaluation of chatgpt on mathematical word problems (mwp)
Paulo Shakarian, Abhinav Koyyalamudi, Noel Ngu, and Lakshmivihari Mareedu. 2023 · 2023
Earlier work this paper cites.
Large language models can be easily distracted by irrelevant context
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed Chi, Nathanael Schärli, and Denny Zhou. 2023 · 2023
Earlier work this paper cites.
Do llms exhibit human-like response biases? a case study in survey design
Lindia Tjuatja, Valerie Chen, Sherry Tongshuang Wu, Ameet Talwalkar, and Graham Neubig. 2023 · 2023
Earlier work this paper cites.
Gpt can solve mathematical problems without a calculator
Zhen Yang, Ming Ding, Qingsong Lv, Zhihuan Jiang, Zehai He, Yuyi Guo, Jinfeng Bai, and Jie Tang. 2023 · 2023
Cited alongside, same era.
How well do large language models perform in arithmetic tasks?
Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, and Songfang Huang. 2023 · 2023
Cited alongside, same era.
Large language models for mathematical reasoning: Progresses and challenges
Janice Ahn, Rishu Verma, Renze Lou, Di Liu, Rui Zhang, and Wenpeng Yin. 2024 · 2024
Cited alongside, same era.
Mathify: Evaluating large language models on mathematical problem solving tasks
Avinash Anand, Mohit Gupta, Kritarth Prasad, Navya Singla, Sanjana Sanjeev, Jatin Kumar, Adarsh Raj Shivam, and Rajiv Ratn Shah. 2024 · 2024
Cited alongside, same era.
How numerical precision affects mathematical reasoning capabilities of llms
Arithmetic with language models: From memorization to computation
Davide Maltoni and Matteo Ferrara. 2024 · 2024
Later among the works it cites.
Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models
Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel, Samy Bengio, and Mehrdad Farajtabar. 2024 · 2024
Later among the works it cites.
OpenAI. 2024 · 2024
Later among the works it cites.
Gpt-4o system card
OpenAI. 2024 · 2024
Later among the works it cites.
Qwen. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Guhao Feng, Kai Yang, Yuntian Gu, Xinyue Ai, Shengjie Luo, Jiacheng Sun, Di He, Zhenguo Li, and Liwei Wang. 2024 · 2024
Cited alongside, same era.
Mathematical capabilities of chatgpt
Simon Frieder, Luca Pinchetti, Ryan-Rhys Griffiths, Tommaso Salvatori, Thomas Lukasiewicz, Philipp Petersen, and Julius Berner. 2024 · 2024
Cited alongside, same era.
Learning beyond pattern matching? assaying mathematical understanding in llms
Siyuan Guo, Aniket Didolkar, Nan Rosemary Ke, Anirudh Goyal, Ferenc Huszár, and Bernhard Schölkopf. 2024 · 2024
Cited alongside, same era.
Pengfei Hong, Deepanway Ghosal, Navonil Majumder, Somak Aditya, Rada Mihalcea, and Soujanya Poria. 2024 · 2024
Cited alongside, same era.
Do large code models understand programming concepts? a black-box approach
Ashish Hooda, Mihai Christodorescu, Miltiadis Allamanis, Aaron Wilson, Kassem Fawaz, and Somesh Jha. 2024 · 2024
Cited alongside, same era.
A peek into token bias: Large language models are not yet genuine reasoners
Bowen Jiang, Yangxinyu Xie, Zhuoqun Hao, Xiaomeng Wang, Tanwi Mallick, Weijie J. Su, Camillo J. Taylor, and Dan Roth. 2024 · 2024
Cited alongside, same era.
Kuei-Chun Kao, Ruochen Wang, and Cho-Jui Hsieh. 2024 · 2024
Cited alongside, same era.
Investigating implicit bias in large language models: A large-scale study of over 50 llms
Ananya Kumar and Siddharth Jain. 2024 · 2024
Cited alongside, same era.
Roy Xie, Chengxuan Huang, Junlin Wang, and Bhuwan Dhingra. 2024 · 2024
Later among the works it cites.
Mr-gsm8k: A meta-reasoning benchmark for large language model evaluation
Zhongshen Zeng, Pengguang Chen, Shu Liu, Haiyun Jiang, and Jiaya Jia. 2024 · 2024
Later among the works it cites.
Arithmattack: Evaluating robustness of llms to noisy context in math problem solving
Zain Ul Abedin, Shahzeb Qamar, Lucie Flek, and Akbar Karimi. 2025 · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. 2025 · 2025
Closest in time.
Introducing OpenAI o1
OpenAI. 2025 · 2025
Closest in time.
Are NLP models really able to solve simple math word problems?
Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021 · 2094
Closest in time.