Fetching the paper…
Reading the bibliography…
Large language models (LLMs) can solve arithmetic word problems with high accuracy, but little is known about how well they generalize to more complex problems.
Untersuchungen über das logische schließen. I
Gerhard Gentzen · 1935
Earlier work this paper cites.
The relative efficiency of propositional proof systems
Stephen A. Cook and Robert A. Reckhow · 1979
Earlier work this paper cites.
The development of semantic categories for addition and subtraction
Pearla Nesher, James G. Greeno, and Mary S. Riley · 1982
Earlier work this paper cites.
Development of Children’s Problem-Solving Ability in Arithmetic , pp. 153–196
Mary Riley, James Greeno, and Joan Heller · 1983
Earlier work this paper cites.
Linear logic
Jean-Yves Girard · 1987
Earlier work this paper cites.
Bootstrap methods: Another look at the jackknife
Bradley Efron · 1992
Earlier work this paper cites.
What makes certain arithmetic word problems involving the comparison of sets so difficult for children?
Elsbeth Stern · 1993
Earlier work this paper cites.
Natural deduction
Frank Pfenning · 2004
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner · 2021
Earlier work this paper cites.
Dynabench: Rethinking benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams · 2021
Earlier work this paper cites.
Are NLP models really able to solve simple math word problems?
Arkil Patel, Satwik Bhattamishra, and Navin Goyal · 2021
Earlier work this paper cites.
Can you learn an algorithm? Generalizing from easy to hard problems with recurrent networks
Avi Schwarzschild, Eitan Borgnia, Arjun Gupta, Furong Huang, Uzi Vishkin, Micah Goldblum, and Tom Goldstein · 2021
Earlier work this paper cites.
Exploring length generalization in large language models
Cem Anil, Yuhuai Wu, Anders Johan Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay Venkatesh Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, and Behnam Neyshabur · 2022
Earlier work this paper cites.
Measuring progress on scalable oversight for large language models, 2022
Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamilė Lukošiūtė, Amanda Askell, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Christopher Olah, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Jackson Kernion, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Liane Lovitt, Nelson Elhage, Nicholas Schiefer, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Robin Larson, Sam McCandlish, Sandipan Kundu, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Ben Mann, and Jared Kaplan · 2022
Earlier work this paper cites.
Curriculum prompt learning with self-training for abstractive dialogue summarization
Changqun Li, Linlin Wang, Xin Lin, Gerard de Melo, and Liang He · 2022
Earlier work this paper cites.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp · 2022
Earlier work this paper cites.
Impact of pretraining term frequencies on few-shot numerical reasoning
Yasaman Razeghi, Robert L Logan IV, Matt Gardner, and Sameer Singh · 2022
Earlier work this paper cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou · 2022
Earlier work this paper cites.
How do in-context examples affect compositional generalization?
Shengnan An, Zeqi Lin, Qiang Fu, Bei Chen, Nanning Zheng, Jian-Guang Lou, and Dongmei Zhang · 2023
Earlier work this paper cites.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision, 2023
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, and Jeff Wu · 2023
Cited alongside, same era.
Faith and fate: Limits of transformers on compositionality
Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jiang, Bill Yuchen Lin, Sean Welleck, Peter West, Chandra Bhagavatula, Ronan Le Bras, Jena D. Hwang, Soumya Sanyal, Xiang Ren, Allyson Ettinger, Zaid Harchaoui, and Yejin Choi · 2023
Cited alongside, same era.
Complexity-based prompting for multi-step reasoning
Yao Fu, Hao Peng, Ashish Sabharwal, Peter Clark, and Tushar Khot · 2023
Cited alongside, same era.
Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks
Alon Jacovi, Avi Caciularu, Omer Goldman, and Yoav Goldberg · 2023
Cited alongside, same era.
CLadder: A benchmark to assess causal reasoning capabilities of language models
A systematic comparison of syllogistic reasoning in humans and language models
Tiwalayo Eisape, Michael Henry Tessler, Ishita Dasgupta, Fei Sha, Sjoerd van Steenkiste, and Tal Linzen · 2024
Closest in time.
What’s in my big data?
Yanai Elazar, Akshita Bhagia, Ian Helgi Magnusson, Abhilasha Ravichander, Dustin Schwenk, Alane Suhr, Evan Pete Walsh, Dirk Groeneveld, Luca Soldaini, Sameer Singh, Hannaneh Hajishirzi, Noah A. Smith, and Jesse Dodge · 2024
Closest in time.
The unreasonable effectiveness of easy training data for hard tasks
Peter Hase, Mohit Bansal, Peter Clark, and Sarah Wiegreffe · 2024
Closest in time.
Case-based or rule-based: How do transformers do the math?
Yi Hu, Xiaojuan Tang, Haotong Yang, and Muhan Zhang · 2024
Closest in time.
MUSTARD: Mastering uniform synthesis of theorem and proof data
Yinya Huang, Xiaohan Lin, Zhengying Liu, Qingxing Cao, Huajian Xin, Haiming Wang, Zhenguo Li, Linqi Song, and Xiaodan Liang · 2024
Closest in time.
A peek into token bias: Large language models are not yet genuine reasoners
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhijing Jin, Yuen Chen, Felix Leeb, Luigi Gresele, Ojasv Kamal, Zhiheng LYU, Kevin Blin, Fernando Gonzalez Adauto, Max Kleiman-Weiner, Mrinmaya Sachan, and Bernhard Schölkopf · 2023
Cited alongside, same era.
Do deep neural networks capture compositionality in arithmetic reasoning?
Keito Kudo, Yoichi Aoki, Tatsuki Kuribayashi, Ana Brassard, Masashi Yoshikawa, Keisuke Sakaguchi, and Kentaro Inui · 2023
Cited alongside, same era.
Transformers learn shortcuts to automata
Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang · 2023
Cited alongside, same era.
World models for math story problems
Andreas Opedal, Niklas Stoehr, Abulhair Saparov, and Mrinmaya Sachan · 2023
Cited alongside, same era.
NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark
Oscar Sainz, Jon Campos, Iker García-Ferrero, Julen Etxaniz, Oier Lopez de Lacalle, and Eneko Agirre · 2023
Cited alongside, same era.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought
Abulhair Saparov and He He · 2023
Cited alongside, same era.
Testing the general deductive reasoning capacity of large language models using OOD examples
Abulhair Saparov, Richard Yuanzhe Pang, Vishakh Padmakumar, Nitish Joshi, Mehran Kazemi, Najoung Kim, and He He · 2023
Cited alongside, same era.
An independent evaluation of ChatGPT on mathematical word problems (MWP)
Paulo Shakarian, Abhinav Koyyalamudi, Noel Ngu, and Lakshmivihari Mareedu · 2023
Cited alongside, same era.
Bowen Jiang, Yangxinyu Xie, Zhuoqun Hao, Xiaomeng Wang, Tanwi Mallick, Weijie J Su, Camillo Jose Taylor, and Dan Roth · 2024
Closest in time.
Lost in the middle: How language models use long contexts
Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang · 2024
Closest in time.
The llama 3 herd of models, 2024
Llama Team · 2024
Closest in time.
Rule extrapolation in language modeling: A study of compositional generalization on OOD prompts
Anna Mészáros, Szilvia Ujváry, Wieland Brendel, Patrik Reizinger, and Ferenc Huszár · 2024
Closest in time.
Mathcamps: Fine-grained synthesis of mathematical problems from human curricula, 2024
Shubhra Mishra, Gabriel Poesia, Belinda Mo, and Noah D. Goodman · 2024
Closest in time.
Do language models exhibit the same cognitive biases in problem solving as human learners?
Andreas Opedal, Alessandro Stolfo, Haruki Shirakami, Ying Jiao, Ryan Cotterell, Bernhard Schölkopf, Abulhair Saparov, and Mrinmaya Sachan · 2024
Closest in time.
OpenAI · 2024
Closest in time.
Deciphering the factors influencing the efficacy of chain-of-thought: Probability, memorization, and noisy reasoning
Akshara Prabhakar, Thomas L. Griffiths, and R. Thomas McCoy · 2024
Closest in time.
Easy-to-hard generalization: Scalable alignment beyond human supervision
Zhiqing Sun, Longhui Yu, Yikang Shen, Weiyang Liu, Yiming Yang, Sean Welleck, and Chuang Gan · 2024
Closest in time.
Limits of transformer language models on learning to compose algorithms
Jonathan Thomm, Giacomo Camposampiero, Aleksandar Terzic, Michael Hersche, Bernhard Schölkopf, and Abbas Rahimi · 2024
Closest in time.
Physics of language models: Part 2.1, grade-school math and the hidden reasoning process, 2024
Tian Ye, Zicheng Xu, Yuanzhi Li, and Zeyuan Allen-Zhu · 2024
Closest in time.
Metamath: Bootstrap your own mathematical questions for large language models
Longhui Yu, Weisen Jiang, Han Shi, Jincheng YU, Zhengying Liu, Yu Zhang, James Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu · 2024
Closest in time.
A careful examination of large language model performance on grade school arithmetic
Hugh Zhang, Jeff Da, Dean Lee, Vaughn Robinson, Catherine Wu, William Song, Tiffany Zhao, Pranav Vishnu Raja, Charlotte Zhuang, Dylan Z Slack, Qin Lyu, Sean M. Hendryx, Russell Kaplan, Michele Lunati, and Summer Yue · 2024
Closest in time.
Deepseek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning, 2025
DeepSeek-AI · 2025
Closest in time.
GSM-symbolic: Understanding the limitations of mathematical reasoning in large language models
Seyed Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel, Samy Bengio, and Mehrdad Farajtabar · 2025
Closest in time.