Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have demonstrated impressive performance on reasoning tasks, including mathematical reasoning.
Children’s reactions to verbal arithmetical problems with missing, surplus or contradictory data
Ewa Puchalska and Zbigniew Semadeni. 1987 · 1987
Earlier work this paper cites.
Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in nlp
John X Morris, Eli Lifland, Jin Yong Yoo, Jake Grigsby, Di Jin, and Yanjun Qi. 2020 · 2005
Earlier work this paper cites.
Z3: an efficient SMT solver
Leonardo Mendonça de Moura and Nikolaj S. Bjørner. 2008 · 2008
Earlier work this paper cites.
The smt-lib standard: Version 2.0
Clark Barrett, Aaron Stump, Cesare Tinelli, et al. 2010 · 2010
Earlier work this paper cites.
Mcts based on simple regret
David Tolpin and Solomon Shimony. 2012 · 2012
Earlier work this paper cites.
Learning to solve arithmetic word problems with verb categorization
Mohammad Javad Hosseini, Hannaneh Hajishirzi, Oren Etzioni, and Nate Kushman. 2014 · 2014
Earlier work this paper cites.
Mawps: A math word problem repository
Rik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, and Hannaneh Hajishirzi. 2016 · 2016
Earlier work this paper cites.
Data imbalance in classification: Experimental evaluation
Fadi Thabtah, Suhel Hammoud, Firuz Kamalov, and Amanda Gonsalves. 2020 · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021 · 2021
Earlier work this paper cites.
Are nlp models really able to solve simple math word problems?
Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021 · 2021
Earlier work this paper cites.
Adversarial glue: A multi-task benchmark for robustness evaluation of language models
Boxin Wang, Chejian Xu, Shuohang Wang, Zhe Gan, Yu Cheng, Jianfeng Gao, Ahmed Hassan Awadallah, and Bo Li. 2021 · 2021
Earlier work this paper cites.
Ai in finance: challenges, techniques, and opportunities
Longbing Cao. 2022 · 2022
Earlier work this paper cites.
Solving quantitative reasoning problems with language models
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. 2022 · 2022
Earlier work this paper cites.
Robust training under label noise by over-parameterization
Sheng Liu, Zhihui Zhu, Qing Qu, and Chong You. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. 2022 · 2022
Cited alongside, same era.
Pal: Program-aided language models
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023 · 2023
Cited alongside, same era.
Visual programming: Compositional visual reasoning without training
Tanmay Gupta and Aniruddha Kembhavi. 2023 · 2023
Cited alongside, same era.
Query and response augmentation cannot help out-of-domain math reasoning generalization
Chengpeng Li, Zheng Yuan, Guanting Dong, Keming Lu, Jiancan Wu, Chuanqi Tan, Xiang Wang, and Chang Zhou. 2023 · 2023
Cited alongside, same era.
Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct
Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Jianguang Lou, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, and Dongmei Zhang. 2023 · 2023
Hui Huang, Yingqi Qu, Jing Liu, Muyun Yang, and Tiejun Zhao. 2024 · 2024
Closest in time.
Qwen2. 5-coder technical report
Binyuan Hui, Jian Yang, Zeyu Cui, Jiaxi Yang, Dayiheng Liu, Lei Zhang, Tianyu Liu, Jiajun Zhang, Bowen Yu, Keming Lu, et al. 2024 · 2024
Closest in time.
Large language models are unconscious of unreasonability in math problems
Jingyuan Ma, Damai Dai, Lei Sha, and Zhifang Sui. 2024 · 2024
Closest in time.
Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models
Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel, Samy Bengio, and Mehrdad Farajtabar. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Logic-lm: Empowering large language models with symbolic solvers for faithful logical reasoning
Liangming Pan, Alon Albalak, Xinyi Wang, and William Yang Wang. 2023 · 2023
Cited alongside, same era.
Large language models can be easily distracted by irrelevant context
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed Chi, Nathanael Schärli, and Denny Zhou. 2023 · 2023
Cited alongside, same era.
Yuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig, and Pengfei Liu. 2023 · 2023
Cited alongside, same era.
Metamath: Bootstrap your own mathematical questions for large language models
Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. 2023 · 2023
Cited alongside, same era.
Large language models as commonsense knowledge for large-scale task planning
Zirui Zhao, Wee Sun Lee, and David Hsu. 2023 · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023 · 2023
Cited alongside, same era.
Large language models for mathematical reasoning: Progresses and challenges
Janice Ahn, Rishu Verma, Renze Lou, Di Liu, Rui Zhang, and Wenpeng Yin. 2024 · 2024
Cited alongside, same era.
Shiyu Tian, Hongxin Wei, Yiqun Wang, and Lei Feng. 2024 · 2024
Closest in time.
Embedding trajectory for out-of-distribution detection in mathematical reasoning
Yiming Wang, Pei Zhang, Baosong Yang, Derek Wong, Zhuosheng Zhang, and Rui Wang. 2024 · 2024
Closest in time.
Satlm: Satisfiability-aided language models using declarative prompting
Xi Ye, Qiaochu Chen, Isil Dillig, and Greg Durrett. 2024 · 2024
Closest in time.
Chatglm: A family of large language models from glm-130b to glm-4 all tools
Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Dan Zhang, Diego Rojas, Guanyu Feng, Hanlin Zhao, et al. 2024 · 2024
Closest in time.
Exploring the compositional deficiency of large language models in mathematical reasoning through trap problems
Jun Zhao, Jingqi Tong, Yurong Mou, Ming Zhang, Qi Zhang, and Xuan-Jing Huang. 2024 · 2024
Closest in time.
rstar-math: Small llms can master math reasoning with self-evolved deep thinking
Xinyu Guan, Li Lyna Zhang, Yifei Liu, Ning Shang, Youran Sun, Yi Zhu, Fan Yang, and Mao Yang. 2025 · 2025
Closest in time.
Mohammad Raza and Natasa Milic-Frayling. 2025 · 2025
Closest in time.
Viscra: A visual chain reasoning attack for jailbreaking multimodal large language models
Bingrui Sima, Linhua Cong, Wenxuan Wang, and Kun He. 2025 · 2025
Closest in time.
Feng Xiong, Hongling Xu, Yifei Wang, Runxi Cheng, Yong Wang, and Xiangxiang Chu. 2025 · 2025
Closest in time.