Fetching the paper…
Reading the bibliography…
Chain-of-Thought (CoT) prompting plays an indispensable role in endowing large language models (LLMs) with complex reasoning capabilities.
Causality
Judea Pearl · 2009
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge, 2019
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2019
Earlier work this paper cites.
Explaining machine learning classifiers through diverse counterfactual explanations
Ramaravind K Mothilal, Amit Sharma, and Chenhao Tan · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Earlier work this paper cites.
Explaining black-box algorithms using probabilistic contrastive counterfactuals
Sainyam Galhotra, Romila Pradhan, and Babak Salimi · 2021
Earlier work this paper cites.
Local explanations via necessity and sufficiency: Unifying theory and practice
David S Watson, Limor Gultchin, Ankur Taly, and Luciano Floridi · 2021
Earlier work this paper cites.
Causal explanations and xai
Sander Beckers · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao · 2022
Earlier work this paper cites.
On the (complete) reasons behind decisions
Adnan Darwiche and Auguste Hirth · 2023
Earlier work this paper cites.
Active prompting with chain-of-thought for large language models
Shizhe Diao, Pengcheng Wang, Yong Lin, Rui Pan, Xiang Liu, and Tong Zhang · 2023
Earlier work this paper cites.
Chessgpt: Bridging policy learning and language modeling
Xidong Feng, Yicheng Luo, Ziyan Wang, Hongrui Tang, Mengyue Yang, Kun Shao, David Mguni, Yali Du, and Jun Wang · 2023
Earlier work this paper cites.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al · 2023
Earlier work this paper cites.
Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman · 2023
Earlier work this paper cites.
Invariant learning via probability of sufficient and necessary causes
Mengyue Yang, Zhen Fang, Yonggang Zhang, Yali Du, Furui Liu, Jean-Francois Ton, Jianhong Wang, and Jun Wang · 2023
Earlier work this paper cites.
Llama 3 model card, 2024
AI@Meta · 2024
Earlier work this paper cites.
Mitigating copy bias in in-context learning through neuron pruning
Ameen Ali, Lior Wolf, and Ivan Titov · 2024
Earlier work this paper cites.
Increasing the llm accuracy for question answering: Ontologies to the rescue!
Dean Allemang and Juan Sequeda · 2024
Earlier work this paper cites.
Cause and effect: can large language models truly understand causality?
Swagata Ashwani, Kshiteesh Hegde, Nishith Reddy Mannuru, Dushyant Singh Sengar, Mayank Jindal, Krishna Chaitanya Rao Kathala, Dishant Banga, Vinija Jain, and Aman Chadha · 2024
Earlier work this paper cites.
Graph of thoughts: Solving elaborate problems with large language models
Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, et al · 2024
Earlier work this paper cites.
Do not think that much for 2+ 3=? on the overthinking of o1-like llms
Xingyu Chen, Jiahao Xu, Tian Liang, Zhiwei He, Jianhui Pang, Dian Yu, Linfeng Song, Qiuzhi Liu, Mengfei Zhou, Zhuosheng Zhang, et al · 2024
Cited alongside, same era.
Compressed chain of thought: Efficient reasoning through dense representations
Jeffrey Cheng and Benjamin Van Durme · 2024
Cited alongside, same era.
Rational metareasoning for large language models
C Nicolò De Sabbata, Theodore R Sumers, and Thomas L Griffiths · 2024
Cited alongside, same era.
Deepseek-v3 technical report, 2024
DeepSeek-AI · 2024
Cited alongside, same era.
Break the chain: Large language models can be shortcut reasoners
Computational experiments for complex social systems: Integrated design of experiment system
Xiao Xue, Xiangning Yu, Deyu Zhou, Xiao Wang, Chongke Bi, Shufang Wang, and Fei-Yue Wang · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan · 2024
Later among the works it cites.
Beyond traditional metrics: The power of value entropy in multidimensional evaluation of the service ecosystem
Xiangning Yu, Xiao Xue, Deyu Zhou, Li Fang, and Zhiyong Feng · 2024
Later among the works it cites.
Dots: Learning to reason dynamically in llms via optimal reasoning trajectories search
Murong Yue, Wenlin Yao, Haitao Mi, Dian Yu, Ziyu Yao, and Dong Yu · 2024
Later among the works it cites.
Empowering llms with logical reasoning: A comprehensive survey
Fengxiang Cheng, Haoxuan Li, Fenrong Liu, Robert Van Rooij, Kun Zhang, and Zhouchen Lin · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mengru Ding, Hanmeng Liu, Zhizhang Fu, Jian Song, Wenbo Xie, and Yue Zhang · 2024
Cited alongside, same era.
Token-budget-aware llm reasoning
Tingxu Han, Chunrong Fang, Shiyu Zhao, Shiqing Ma, Zhenyu Chen, and Zhenting Wang · 2024
Cited alongside, same era.
Training large language models to reason in a continuous latent space
Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, and Yuandong Tian · 2024
Cited alongside, same era.
Reasoning elicitation in language models via counterfactual feedback
Alihan Hüyük, Xinnuo Xu, Jacqueline Maasch, Aditya V Nori, and Javier González · 2024
Cited alongside, same era.
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al · 2024
Cited alongside, same era.
Self-harmonized chain of thought
Ziqi Jin and Wei Lu · 2024
Cited alongside, same era.
C3ot: Generating shorter chain-of-thought without compromising effectiveness
Yu Kang, Xianghui Sun, Liangyu Chen, and Wei Zou · 2024
Cited alongside, same era.
Openai-o1 ab testing: Does the o1 model really do good reasoning in math problem solving?
Leo Li, Ye Luo, and Tingyou Pan · 2024
Cited alongside, same era.
Closest in time.
Aime 2025 dataset
OpenCompass Contributors · 2025
Closest in time.
The danger of overthinking: Examining the reasoning-action dilemma in agentic tasks
Alejandro Cuadron, Dacheng Li, Wenjie Ma, Xingyao Wang, Yichuan Wang, Siyuan Zhuang, Shu Liu, Luis Gaspar Schroeder, Tian Xia, Huanzhi Mao, et al · 2025
Closest in time.
Yingqian Cui, Pengfei He, Jingying Zeng, Hui Liu, Xianfeng Tang, Zhenwei Dai, Yan Han, Chen Luo, Jing Huang, Zhen Li, et al · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI · 2025
Closest in time.
Large language models are demonstration pre-selectors for themselves
Jiarui Jin, Yuwei Wu, Haoxuan Li, Xiaoting He, Weinan Zhang, Yiming Yang, Yong Yu, Jun Wang, and Mengyue Yang · 2025
Closest in time.
Zebralogic: On the scaling limits of llms for logical reasoning
Bill Yuchen Lin, Ronan Le Bras, Kyle Richardson, Ashish Sabharwal, Radha Poovendran, Peter Clark, and Yejin Choi · 2025
Closest in time.
Thought manipulation: External thought can be efficient for large reasoning models
Yule Liu, Jingyi Zheng, Zhen Sun, Zifan Peng, Wenhan Dong, Zeyang Sha, Shiwen Cui, Weiqiang Wang, and Xinlei He · 2025
Closest in time.
Self-training elicits concise reasoning in large language models
Tergel Munkhbat, Namgyu Ho, Seohyun Kim, Yongjin Yang, Yujin Kim, and Se-Young Yun · 2025
Closest in time.
Learning to reason with llms, 2025
OpenAI · 2025
Closest in time.
Qwq: Reflect deeply on the boundaries of the unknown
Team Qwen · 2025
Closest in time.
Trust me, i’m wrong: High-certainty hallucinations in llms
Adi Simhi, Itay Itzhak, Fazl Barez, Gabriel Stanovsky, and Yonatan Belinkov · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, et al · 2025
Closest in time.
Effectively controlling reasoning models through thinking intervention
Tong Wu, Chong Xiang, Jiachen T Wang, and Prateek Mittal · 2025
Closest in time.
Towards system 2 reasoning in llms: Learning how to think with meta chain-of-though
Violet Xiang, Charlie Snell, Kanishk Gandhi, Alon Albalak, Anikait Singh, Chase Blagden, Duy Phung, Rafael Rafailov, Nathan Lile, Dakota Mahan, et al · 2025
Closest in time.
Towards thinking-optimal scaling of test-time compute for llm reasoning
Wenkai Yang, Shuming Ma, Yankai Lin, and Furu Wei · 2025
Closest in time.