Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have demonstrated remarkable mathematical capabilities, largely driven by chain-of-thought (CoT) prompting, which decomposes complex reasoning into step-by-step solutions.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena · 2021
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Dual instruction tuning with large language models for mathematical reasoning
Yongwei Zhou and Tiejun Zhao · 2022
Earlier work this paper cites.
Neural discovery of permutation subgroups
Pavan Karjol, Rohan Kashyap, and AP Prathosh · 2023
Earlier work this paper cites.
Goat: Fine-tuned llama outperforms gpt-4 on arithmetic tasks
Tiedong Liu and Bryan Kian Hsiang Low · 2023
Earlier work this paper cites.
Improving large language model fine-tuning for solving math problems
Yixin Liu, Avi Singh, C. Daniel Freeman, John D. Co-Reyes, and Peter J. Liu · 2023
Earlier work this paper cites.
Auto-regressive next-token predictors are universal learners
Eran Malach · 2023
Earlier work this paper cites.
Alessandro Stolfo, Yonatan Belinkov, and Mrinmaya Sachan · 2023
Earlier work this paper cites.
Latent space symmetry discovery
Jianke Yang, Nima Dehmamy, Robin Walters, and Rose Yu · 2023
Cited alongside, same era.
Claude 3 haiku: our fastest model yet
Anthropic · 2024
Cited alongside, same era.
Language models are symbolic learners in arithmetic
Chunyuan Deng, Zhiqi Li, Roy Xie, Ruidi Chang, and Hanjie Chen · 2024
Cited alongside, same era.
Language models know the value of numbers, 2024
Zhifang Sui Fangwei Zhu, Damai Dai · 2024
Cited alongside, same era.
Towards revealing the mystery behind chain of thought: a theoretical perspective
Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang · 2024
Cited alongside, same era.
Learning mathematical rules with large language models, 2024
Language models encode numbers using digit representations in base 10
Amit Arnold Levy and Mor Geva · 2024
Closest in time.
Bohan Lyu, Yadi Cao, Duncan Watson-Parris, Leon Bergen, Taylor Berg-Kirkpatrick, and Rose Yu · 2024
Closest in time.
Why think step by step? reasoning emerges from the locality of experience
Ben Prystawski, Michael Li, and Noah Goodman · 2024
Closest in time.
Numerologic: Number encoding for enhanced llms’ numerical reasoning
Eli Schwartz, Leshem Choshen, Joseph Shtok, Sivan Doveh, Leonid Karlinsky, and Assaf Arbelle · 2024
Closest in time.
Revorder: A novel method for enhanced arithmetic in language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Antoine Gorceix, Bastien Le Chenadec, Ahmad Rammal, Nelson Vadori, and Manuela Veloso · 2024
Cited alongside, same era.
Exploring reversal mathematical reasoning ability for large language models
Pei Guo, WangJie You, Juntao Li, Yan Bowen, and Min Zhang · 2024
Cited alongside, same era.
How does gpt-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model
Michael Hanna, Ollie Liu, and Alexandre Variengien · 2024
Cited alongside, same era.
Exploring group and symmetry principles in large language models
Shima Imani and Hamid Palangi · 2024
Cited alongside, same era.
A unified framework for discovering discrete symmetries
Pavan Karjol, Rohan Kashyap, Aditya Gopalan, and AP Prathosh · 2024
Cited alongside, same era.
Executing arithmetic: Fine-tuning large language models as turing machines
Junyu Lai, Jiahe Xu, Yao Yang, Yunpeng Huang, Chun Cao, and Jingwei Xu · 2024
Cited alongside, same era.
Si Shen, Peijun Shen, and Danhao Zhu · 2024
Closest in time.
Mathscale: Scaling instruction tuning for mathematical reasoning
Zhengyang Tang, Xingxing Zhang, Benyou Wang, and Furu Wei · 2024
Closest in time.
A theory for length generalization in learning to reason
Changnan Xiao and Bing Liu · 2024
Closest in time.
Scaffolding learning: From specific to generic with large language models
David S. Yin and Xiaoxin Yin · 2024
Closest in time.
Interpreting and improving large language models in arithmetic calculation
Wei Zhang, Wan Chaoqun, Yonggang Zhang, Yiu Ming Cheung, Xinmei Tian, Xu Shen, and Jieping Ye · 2024
Closest in time.
Arithmeticgpt: Empowering small-size large language models with advanced arithmetic skills
Zitao Liu, Ying Zheng, Zhibo Yin, Jiahao Chen, Tianqiao Liu, Mi Tian, and Weiqi Luo · 2025
Closest in time.