Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) exhibit impressive reasoning abilities, yet their reliance on structured step-by-step processing reveals a critical limitation.
Hedges: A study in meaning criteria and the logic of fuzzy concepts
George Lakoff · 1973
Earlier work this paper cites.
Judgment under uncertainty: Heuristics and biases: Biases in judgments reveal some heuristics of thinking under uncertainty
Amos Tversky and Daniel Kahneman · 1974
Earlier work this paper cites.
Goal-directed instrumental action: contingency and incentive learning and their cortical substrates
Bernard W Balleine and Anthony Dickinson · 1998
Earlier work this paper cites.
Advancing the rationality debate
Keith E Stanovich and Richard F West · 2000
Earlier work this paper cites.
Uncertainty-based competition between prefrontal and dorsolateral striatal systems for behavioral control
Nathaniel D Daw, Yael Niv, and Peter Dayan · 2005
Earlier work this paper cites.
Model-based influences on humans’ choices and striatal prediction errors
Nathaniel D Daw, Samuel J Gershman, Ben Seymour, Peter Dayan, and Raymond J Dolan · 2011
Earlier work this paper cites.
Thinking, fast and slow
Daniel Kahneman · 2011
Earlier work this paper cites.
Speed/accuracy trade-off between the habitual and the goal-directed processes
Mehdi Keramati, Amir Dezfouli, and Payam Piray · 2011
Earlier work this paper cites.
Goals and habits in the brain
Ray J Dolan and Peter Dayan · 2013
Earlier work this paper cites.
Dual-process theories of higher cognition: Advancing the debate
Jonathan St BT Evans and Keith E Stanovich · 2013
Earlier work this paper cites.
Learning to solve arithmetic word problems with verb categorization
Mohammad Javad Hosseini, Hannaneh Hajishirzi, Oren Etzioni, and Nate Kushman · 2014
Earlier work this paper cites.
Neural computations underlying arbitration between model-based and model-free learning
Sang Wan Lee, Shinsuke Shimojo, and John P O’doherty · 2014
Earlier work this paper cites.
Parsing algebraic word problems into equations
Rik Koncel-Kedziorski, Hannaneh Hajishirzi, Ashish Sabharwal, Oren Etzioni, and Siena Dumas Ang · 2015
Earlier work this paper cites.
Solving general arithmetic word problems
Subhro Roy and Dan Roth · 2015
Earlier work this paper cites.
Characterizing a psychiatric symptom dimension related to deficits in goal-directed control
Claire M Gillan, Michal Kosinski, Robert Whelan, Elizabeth A Phelps, and Nathaniel D Daw · 2016
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom · 2017
Earlier work this paper cites.
Dorsal hippocampus contributes to model-based planning
Kevin J Miller, Matthew M Botvinick, and Carlos D Brody · 2017
Earlier work this paper cites.
Prioritized memory access explains planning and hippocampal replay
Marcelo G Mattar and Nathaniel D Daw · 2018
Earlier work this paper cites.
Hedging, weasel words, and truthiness in scientific writing
Douglas E Ott · 2018
Earlier work this paper cites.
Socialiqa: Commonsense reasoning about social interactions
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras, and Yejin Choi · 2019
Earlier work this paper cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Dissociating neural learning signals in human sign-and goal-trackers
Daniel J Schad, Michael A Rapp, Maria Garbusow, Stephan Nebe, Miriam Sebold, Elisabeth Obst, Christian Sommer, Lorenz Deserno, Milena Rabovsky, Eva Friedel, et al · 2020
Earlier work this paper cites.
Thinking fast and slow in ai
Grady Booch, Francesco Fabiano, Lior Horesh, Kiran Kate, Jonathan Lenchner, Nick Linck, Andreas Loreggia, Keerthiram Murgesan, Nicholas Mattei, Francesca Rossi, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies
Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
Are NLP models really able to solve simple math word problems?
Arkil Patel, Satwik Bhattamishra, and Navin Goyal · 2021
Earlier work this paper cites.
Linear reinforcement learning in planning, grid fields, and cognitive control
Payam Piray and Nathaniel D Daw · 2021
Earlier work this paper cites.
COM2SENSE: A commonsense reasoning benchmark with complementary sentences
Shikhar Singh, Nuan Wen, Yu Hou, Pegah Alipoormolabashi, Te-lin Wu, Xuezhe Ma, and Nanyun Peng · 2021
Earlier work this paper cites.
Career decision-making from a dual-process perspective: Looking back, looking forward
Hui Xu · 2021
Cited alongside, same era.
Bounded reflectivism and epistemic identity
Nick Byrd · 2022
Cited alongside, same era.
Bertopic: Neural topic modeling with a class-based tf-idf procedure
Maarten Grootendorst · 2022
Cited alongside, same era.
System 1 + system 2 = better world: Neural-symbolic chain of logic reasoning
Wenyue Hua and Yongfeng Zhang · 2022
Cited alongside, same era.
Large language models can self-improve
Jiaxin Huang, Shixiang Shane Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han · 2022
Cited alongside, same era.
Cognitive bias in decision-making with llms
Jessica Echterhoff, Yao Liu, Abeer Alessa, Julian McAuley, and Zexue He · 2024
Later among the works it cites.
Thinking fair and slow: On the efficacy of structured prompts for debiasing language models
Shaz Furniturewala, Surgan Jandial, Abhinav Java, Pragyan Banerjee, Simra Shahid, Sumit Bhatia, and Kokil Jaidka · 2024
Later among the works it cites.
Planning like human: A dual-process framework for dialogue planning
Tao He, Lizi Liao, Yixin Cao, Yuanxing Liu, Ming Liu, Zerui Chen, and Bing Qin · 2024
Later among the works it cites.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al · 2024
Later among the works it cites.
A peek into token bias: Large language models are not yet genuine reasoners
Bowen Jiang, Yangxinyu Xie, Zhuoqun Hao, Xiaomeng Wang, Tanwi Mallick, Weijie J Su, Camillo Jose Taylor, and Dan Roth · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jie Huang and Kevin Chen-Chuan Chang · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Cited alongside, same era.
A neural-symbolic approach to natural language understanding
Zhixuan Liu, Zihao Wang, Yuan Lin, and Hang Li · 2022
Cited alongside, same era.
Teaching small language models to reason
Lucie Charlotte Magister, Jonathan Mallinson, Jakub Adamek, Eric Malmi, and Aliaksei Severyn · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Large language models still can’t plan (a benchmark for llms on planning and reasoning about change)
Karthik Valmeekam, Alberto Olmo, Sarath Sreedharan, and Subbarao Kambhampati · 2022
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2022
Cited alongside, same era.
Later among the works it cites.
Mahammed Kamruzzaman and Gene Louis Kim · 2024
Later among the works it cites.
Better zero-shot reasoning with role-play prompting, 2024
Aobo Kong, Shiwan Zhao, Hao Chen, Qicheng Li, Yong Qin, Ruiqi Sun, Xin Zhou, Enzhi Wang, and Xiaohang Dong · 2024
Later among the works it cites.
Language grounded multi-agent reinforcement learning with human-interpretable communication
Huao Li, Hossein Nourkhiz Mahjoub, Behdad Chalaki, Vaishnav Tadiparthi, Kwonjoon Lee, Ehsan Moradi Pari, Charles Lewis, and Katia Sycara · 2024
Later among the works it cites.
Simpo: Simple preference optimization with a reference-free reward
Yu Meng, Mengzhou Xia, and Danqi Chen · 2024
Later among the works it cites.
Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models
Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel, Samy Bengio, and Mehrdad Farajtabar · 2024
Later among the works it cites.
Beyond accuracy: Evaluating the reasoning behavior of large language models–a survey
Philipp Mondorf and Barbara Plank · 2024
Later among the works it cites.
Dynathink: Fast or slow? a dynamic decision-making framework for large language models
Jiabao Pan, Yan Zhang, Chen Zhang, Zuozhu Liu, Hongwei Wang, and Haizhou Li · 2024
Later among the works it cites.
Logicbench: Towards systematic evaluation of logical reasoning ability of large language models
Mihir Parmar, Nisarg Patel, Neeraj Varshney, Mutsumi Nakamura, Man Luo, Santosh Mashetty, Arindam Mitra, and Chitta Baral · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Later among the works it cites.
Arn: Analogical reasoning on narratives
Zhivar Sourati, Filip Ilievski, Pia Sommerauer, and Yifan Jiang · 2024
Later among the works it cites.
Chain-of-thought reasoning without prompting
Xuezhi Wang and Denny Zhou · 2024
Later among the works it cites.
HealMe: Harnessing cognitive reframing in large language models for psychotherapy
Mengxi Xiao, Qianqian Xie, Ziyan Kuang, Zhicheng Liu, Kailai Yang, Min Peng, Weiguang Han, and Jimin Huang · 2024
Later among the works it cites.
Llm2: Let large language models harness system 2 reasoning
Cheng Yang, Chufan Shi, Siheng Li, Bo Shui, Yujiu Yang, and Wai Lam · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan · 2024
Later among the works it cites.
Distilling system 2 into system 1
Ping Yu, Jing Xu, Jason Weston, and Ilia Kulikov · 2024
Later among the works it cites.
Mr-ben: A meta-reasoning benchmark for evaluating system-2 thinking in llms
Zhongshen Zeng, Yinhong Liu, Yingjia Wan, Jingyao Li, Pengguang Chen, Jianbo Dai, Yuxuan Yao, Rongwu Xu, Zehan Qi, Wanru Zhao, et al · 2024
Later among the works it cites.
LLMs learn task heuristics from demonstrations: A heuristic-driven prompting strategy for document-level event argument extraction
Hanzhang Zhou, Junlang Qian, Zijian Feng, Lu Hui, Zixiao Zhu, and Kezhi Mao · 2024
Later among the works it cites.
The danger of overthinking: Examining the reasoning-action dilemma in agentic tasks
Alejandro Cuadron, Dacheng Li, Wenjie Ma, Xingyao Wang, Yichuan Wang, Siyuan Zhuang, Shu Liu, Luis Gaspar Schroeder, Tian Xia, Huanzhi Mao, et al · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Can large language models detect errors in long chain-of-thought reasoning?
Yancheng He, Shilong Li, Jiaheng Liu, Weixun Wang, Xingyuan Bu, Ge Zhang, Zhongyuan Peng, Zhaoxiang Zhang, Zhicheng Zheng, Wenbo Su, et al · 2025
Closest in time.
Wakenllm: Evaluating reasoning potential and stability in llms via fine-grained benchmarking
Zipeng Ling, Yuehao Tang, Shuliang Liu, Junqi Yang, Shenghong Fu, Chen Huang, Kejia Huang, Yao Wan, Zhichao Hou, and Xuming Hu · 2025
Closest in time.
System 1.x: Learning to balance fast and slow planning with language models
Swarnadeep Saha, Archiki Prasad, Justin Chen, Peter Hase, Elias Stengel-Eskin, and Mohit Bansal · 2025
Closest in time.
Parshin Shojaee, Iman Mirzadeh, Keivan Alizadeh, Maxwell Horton, Samy Bengio, and Mehrdad Farajtabar · 2025
Closest in time.
Dualformer: Controllable fast and slow thinking by learning with randomized reasoning traces
DiJia Su, Sainbayar Sukhbaatar, Michael Rabbat, Yuandong Tian, and Qinqing Zheng · 2025
Closest in time.
Probabilistic soundness guarantees in llm reasoning chains
Weiqiu You, Anton Xue, Shreya Havaldar, Delip Rao, Helen Jin, Chris Callison-Burch, and Eric Wong · 2025
Closest in time.