Fetching the paper…
Reading the bibliography…
There is increasing interest in employing large language models (LLMs) as cognitive models.
La naissance de l’intelligence chez l’enfant
Jean Piaget. 1936 · 1936
Earlier work this paper cites.
Verbal cues as an interfering factor in verbal problem solving
Perla Nesher and Eva Teubal. 1975 · 1975
Earlier work this paper cites.
The role of short-term working memory in mental arithmetic
Graham J Hitch. 1978 · 1978
Earlier work this paper cites.
The development of semantic categories for addition and subtraction
P. Nesher, James G. Greeno, and Mary S. Riley. 1982 · 1982
Earlier work this paper cites.
Correspondences and numerical differences between disjoint sets
Tom Hudson. 1983 · 1983
Earlier work this paper cites.
Mental models: towards a cognitive science of language, inference and consciousness , volume 6 of Cognitive Science Series
Philip Nicholas Johnson-Laird. 1983 · 1983
Earlier work this paper cites.
Development of Children’s Problem-Solving Ability in Arithmetic , page 153–196. Learning Research and Development Center, University of Pittsburgh
Mary Riley, James Greeno, and Joan Heller. 1983 · 1983
Earlier work this paper cites.
An integrated model of skill in solving elementary word problems
Diane J. Briars and Jill H. Larkin. 1984 · 1984
Earlier work this paper cites.
The acquisition of addition and subtraction concepts in grades one through three
Thomas P. Carpenter and James M. Moser. 1984 · 1984
Earlier work this paper cites.
Understanding and solving word arithmetic problems
Walter Kintsch and James G. Greeno. 1985 · 1985
Earlier work this paper cites.
Students’ miscomprehension of relational statements in arithmetic word problems
Anne Bovenmyer Lewis and Richard E. Mayer. 1987 · 1987
Earlier work this paper cites.
The role of understanding in solving word problems
Denise Dellarosa Cummins, Walter Kintsch, Kurt Reusser, and Rhonda Weimer. 1988 · 1988
Earlier work this paper cites.
The development of children’s mental multiplication skills
John W. Koshmider and Mark H. Ashcraft. 1991 · 1991
Earlier work this paper cites.
Chapter 8 working memory, automaticity, and problem difficulty
Mark H. Ashcraft, Rick D. Donley, Margaret A. Halas, and Mary Vakali. 1992 · 1992
Earlier work this paper cites.
What makes certain arithmetic word problems involving the comparison of sets so difficult for children?
Elsbeth Stern. 1993 · 1993
Earlier work this paper cites.
Applications of simulated students: An exploration
Kurt VanLehn, Stellan Ohlsson, and Rod Nason. 1994 · 1994
Earlier work this paper cites.
Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing
Yoav Benjamini and Yosef Hochberg. 1995 · 1995
Earlier work this paper cites.
Comprehension of arithmetic word problems: A comparison of successful and unsuccessful problem solvers
Mary Hegarty, Richard E. Mayer, and Christopher A. Monk. 1995 · 1995
Earlier work this paper cites.
Separate roles for executive and phonological components of working memory in mental arithmetic
Ansgar J. Fürst and Graham J. Hitch. 2000 · 2000
Earlier work this paper cites.
Influence of situational and conceptual rewording on word problem solving
Santiago. Vicente, Jose. Orrantia, and Lieven. Verschaffel. 2007 · 2007
Earlier work this paper cites.
Language affects symbolic arithmetic in children: The case of number word inversion
Silke M. Göbel, Korbinian Moeller, Silvia Pixner, Liane Kaufmann, and Hans-Christoph Nuerk. 2014 · 2014
Earlier work this paper cites.
Learning from the folly of others: Learning to self-correct by monitoring the reasoning of virtual characters in a computer-supported mathematics learning environment
Sandra Y. Okita. 2014 · 2014
Earlier work this paper cites.
Word problems: a review of linguistic and numerical factors contributing to their difficulty
Gabriella Daroczy, Magdalena Wolska, Walt Detmar Meurers, and Hans-Christoph Nuerk. 2015 · 2015
Earlier work this paper cites.
Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction
Guido W. Imbens and Donald B. Rubin. 2015 · 2015
Earlier work this paper cites.
Personalized mathematical word problem generation
Oleksandr Polozov, Eleanor O’Rourke, Adam M. Smith, Luke Zettlemoyer, Sumit Gulwani, and Zoran Popovic. 2015 · 2015
Cited alongside, same era.
A theme-rewriting approach for generating algebra word problems
Rik Koncel-Kedziorski, Ioannis Konstas, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2016a · 2016
Cited alongside, same era.
MAWPS: A math word problem repository
Rik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, and Hannaneh Hajishirzi. 2016b · 2016
Cited alongside, same era.
A diverse corpus for evaluating and developing English math word problem solvers
Shen-yun Miao, Chao-Chun Liang, and Keh-Yih Su. 2020 · 2020
Cited alongside, same era.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021 · 2021
Cited alongside, same era.
Automatic educational question generation with difficulty level controls
Ying Jiao, Kumar Shridhar, Peng Cui, Wangchunshu Zhou, and Mrinmaya Sachan. 2023 · 2023
Later among the works it cites.
CLadder: A benchmark to assess causal reasoning capabilities of language models
Zhijing Jin, Yuen Chen, Felix Leeb, Luigi Gresele, Ojasv Kamal, Zhiheng Lyu, Kevin Blin, Fernando Gonzalez Adauto, Max Kleiman-Weiner, Mrinmaya Sachan, and Bernhard Schölkopf. 2023 · 2023
Later among the works it cites.
Beyond English: Evaluating LLMs for Arabic grammatical error correction
Sang Kwon, Gagan Bhatia, El Moatez Billah Nagoudi, and Muhammad Abdul-Mageed. 2023 · 2023
Later among the works it cites.
Simulated learners in educational technology: A systematic literature review and a turing-like test
Tanja Käser and Giora Alexandron. 2023 · 2023
Later among the works it cites.
Exploring effectiveness of GPT-3 in grammatical error correction: A study on performance and controllability in prompt-based methods
Mengsay Loem, Masahiro Kaneko, Sho Takase, and Naoaki Okazaki. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Can you clarify what you said?”: Studying the impact of tutee agents’ follow-up questions on tutors’ learning
Tasmia Shahriar and Noboru Matsuda. 2021 · 2021
Cited alongside, same era.
Language models as agent models
Jacob Andreas. 2022 · 2022
Cited alongside, same era.
Causal inference in natural language processing: Estimation, prediction, interpretation and beyond
Amir Feder, Katherine A. Keith, Emaad Manzoor, Reid Pryzant, Dhanya Sridhar, Zach Wood-Doughty, Jacob Eisenstein, Justin Grimmer, Roi Reichart, Margaret E. Roberts, Brandon M. Stewart, Victor Veitch, and Diyi Yang. 2022 · 2022
Cited alongside, same era.
Capturing failures of large language models via human cognitive biases
Erik Jones and Jacob Steinhardt. 2022 · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Gray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Cited alongside, same era.
Impact of pretraining term frequencies on few-shot numerical reasoning
Yasaman Razeghi, Robert L Logan IV, Matt Gardner, and Sameer Singh. 2022 · 2022
Cited alongside, same era.
Later among the works it cites.
MathDial: A dialogue tutoring dataset with rich pedagogical properties grounded in math reasoning problems
Jakub Macina, Nico Daheim, Sankalan Chowdhury, Tanmay Sinha, Manu Kapur, Iryna Gurevych, and Mrinmaya Sachan. 2023 · 2023
Later among the works it cites.
Dissociating language and thought in large language models: a cognitive perspective
Kyle Mahowald, Anna A Ivanova, Idan A Blank, Nancy Kanwisher, Joshua B Tenenbaum, and Evelina Fedorenko. 2023 · 2023
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt. 2023 · 2023
Later among the works it cites.
Manh Hung Nguyen, Sebastian Tschiatschek, and Adish Singla. 2023 · 2023
Later among the works it cites.
World models for math story problems
Andreas Opedal, Niklas Stoehr, Abulhair Saparov, and Mrinmaya Sachan. 2023 · 2023
Later among the works it cites.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought
Abulhair Saparov and He He. 2023 · 2023
Later among the works it cites.
Ruffle&Riley: Towards the automated induction of conversational tutoring systems
Robin Schmucker, Meng Xia, Amos Azaria, and Tom Mitchell. 2023 · 2023
Later among the works it cites.
Numeric magnitude comparison effects in large language models
Raj Shah, Vijay Marupudi, Reba Koenen, Khushi Bhardwaj, and Sashank Varma. 2023 · 2023
Later among the works it cites.
On second thought, let’s not think step by step! bias and toxicity in zero-shot reasoning
Omar Shaikh, Hongxin Zhang, William Held, Michael Bernstein, and Diyi Yang. 2023 · 2023
Later among the works it cites.
Positional description matters for transformers arithmetic
Ruoqi Shen, Sébastien Bubeck, Ronen Eldan, Yin Tat Lee, Yuanzhi Li, and Yi Zhang. 2023 · 2023
Later among the works it cites.
A mechanistic interpretation of arithmetic reasoning in language models using causal mediation analysis
Alessandro Stolfo, Yonatan Belinkov, and Mrinmaya Sachan. 2023a · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023 · 2023
Later among the works it cites.
A systematic comparison of syllogistic reasoning in humans and language models
Tiwalayo Eisape, MH Tessler, Ishita Dasgupta, Fei Sha, Sjoerd van Steenkiste, and Tal Linzen. 2024 · 2024
Closest in time.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2024 · 2024
Closest in time.
Can large language models infer causation from correlation?
Zhijing Jin, Jiarui Liu, Zhiheng Lyu, Spencer Poff, Mrinmaya Sachan, Rada Mihalcea, Mona Diab, and Bernhard Schölkopf. 2024 · 2024
Closest in time.
OpenAI. 2024 · 2024
Closest in time.
Understanding addition in transformers
Philip Quirke and Fazl Barez. 2024 · 2024
Closest in time.
Large language models as optimizers
Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V Le, Denny Zhou, and Xinyun Chen. 2024 · 2024
Closest in time.
Are NLP models really able to solve simple math word problems?
Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021 · 2094
Closest in time.