Fetching the paper…
Reading the bibliography…
Being prompted to engage in reasoning has emerged as a core technique for using large language models (LLMs), deploying additional inference-time compute to improve task performance.
A systematic review of green ¡scp¿ai¡/scp¿
Roberto Verdecchia, June Sallou, and Luís Cruz · 1942
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan, and Geoffrey E. Hinton · 1991
Earlier work this paper cites.
Principles of metareasoning
Stuart Russell and Eric Wefald · 1991
Earlier work this paper cites.
Rationality and intelligence
Stuart J. Russell · 1997
Earlier work this paper cites.
Measuring massive multitask language understanding, 2021
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2009
Earlier work this paper cites.
Proofwriter: Generating implications, proofs, and abductive statements over natural language, 2021
Oyvind Tafjord, Bhavana Dalvi Mishra, and Peter Clark · 2012
Earlier work this paper cites.
Thinking, fast and slow
Daniel Kahneman · 2014
Earlier work this paper cites.
Thinking fast and slow with deep learning and tree search, 2017
Thomas Anthony, Zheng Tian, and David Barber · 2017
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks, 2017
Alex Graves · 2017
Earlier work this paper cites.
Strategy selection as rational metareasoning
Falk Lieder and Thomas L. Griffiths · 2017
Earlier work this paper cites.
Learning to select computations
Frederick Callaway, Sayan Gul, Paul Krueger, Thomas L. Griffiths, and Falk Lieder · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Rational metareasoning and the plasticity of cognitive control
Falk Lieder, Amitai Shenhav, Sebastian Musslick, and Thomas L. Griffiths · 2018
Earlier work this paper cites.
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser · 2019
Earlier work this paper cites.
Doing more with less: meta-reasoning and meta-learning in humans and machines
Thomas L Griffiths, Frederick Callaway, Michael B Chang, Erin Grant, Paul M Krueger, and Falk Lieder · 2019
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge, 2019
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2019
Earlier work this paper cites.
Understanding human intelligence through human limitations
Thomas L. Griffiths · 2020
Earlier work this paper cites.
Pondernet: Learning to ponder, 2021
Andrea Banino, Jan Balaguer, and Charles Blundell · 2021
Earlier work this paper cites.
Fixation patterns in simple choice reflect optimal information sampling
Frederick Callaway, Antonio Rangel, and Thomas L Griffiths · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
Rational use of cognitive resources in human planning
Frederick Callaway, Bas van Opheusden, Sayan Gul, Priyam Das, Paul M Krueger, Thomas L Griffiths, and Falk Lieder · 2022
Earlier work this paper cites.
Palm: Scaling language modeling with pathways, 2022
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Earlier work this paper cites.
Large language models can self-improve, 2022
Jiaxin Huang, Shixiang Shane Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han · 2022
Earlier work this paper cites.
Time spent thinking in online chess reflects the value of computation
Evan Russek, Daniel Acosta-Kane, Bas van Opheusden, Marcelo G Mattar, and Tom Griffiths · 2022
Cited alongside, same era.
Confident adaptive language modeling, 2022
Tal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani, Dara Bahri, Vinh Q. Tran, Yi Tay, and Donald Metzler · 2022
Cited alongside, same era.
Star: Bootstrapping reasoning with reasoning, 2022
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah D. Goodman · 2022
Cited alongside, same era.
Mixture-of-experts with expert choice routing, 2022
Yanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du, Yanping Huang, Vincent Zhao, Andrew Dai, Zhifeng Chen, Quoc Le, and James Laudon · 2022
Cited alongside, same era.
The growing energy footprint of artificial intelligence
Alex de Vries · 2023
Routellm: Learning to route llms with preference data, 2024
Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M Waleed Kadous, and Ion Stoica · 2024
Closest in time.
Openai o1 system card
OpenAI · 2024
Closest in time.
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, et al · 2024
Closest in time.
Disentangling length from quality in direct preference optimization, 2024
Ryan Park, Rafael Rafailov, Stefano Ermon, and Chelsea Finn · 2024
Closest in time.
Making reasoning matter: Measuring and improving faithfulness of chain-of-thought reasoning, 2024
Debjit Paul, Robert West, Antoine Bosselut, and Boi Faltings · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reinforced self-training (rest) for language modeling, 2023
Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan, Ksenia Konyushkova, Lotte Weerts, Abhishek Sharma, Aditya Siddhant, Alex Ahern, Miaosen Wang, Chenjie Gu, Wolfgang Macherey, Arnaud Doucet, Orhan Firat, and Nando de Freitas · 2023
Cited alongside, same era.
Large language models are zero-shot reasoners, 2023
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2023
Cited alongside, same era.
Fast inference from transformers via speculative decoding, 2023
Yaniv Leviathan, Matan Kalman, and Yossi Matias · 2023
Cited alongside, same era.
Symbolic chain-of-thought distillation: Small models can also “think” step-by-step
Liunian Harold Li, Jack Hessel, Youngjae Yu, Xiang Ren, Kai-Wei Chang, and Yejin Choi · 2023
Cited alongside, same era.
Self-refine: Iterative refinement with self-feedback, 2023
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark · 2023
Cited alongside, same era.
Cotformer: More tokens with attention make up for less depth, 2023
Amirkeivan Mohtashami, Matteo Pagliardini, and Martin Jaggi · 2023
Cited alongside, same era.
Why think step by step? reasoning emerges from the locality of experience, 2023
Ben Prystawski, Michael Y. Li, and Noah D. Goodman · 2023
Cited alongside, same era.
Closest in time.
Let’s think dot by dot: Hidden computation in transformer language models, 2024
Jacob Pfau, William Merrill, and Samuel R. Bowman · 2024
Closest in time.
Synergy-of-thoughts: Eliciting efficient reasoning in hybrid language models, 2024
Yu Shang, Yu Li, Fengli Xu, and Yong Li · 2024
Closest in time.
A long way to go: Investigating length correlations in rlhf, 2024
Prasann Singhal, Tanya Goyal, Jiacheng Xu, and Greg Durrett · 2024
Closest in time.
Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2024
Closest in time.
To cot or not to cot? chain-of-thought helps mainly on math and symbolic reasoning, 2024
Zayne Sprague, Fangcong Yin, Juan Diego Rodriguez, Dongwei Jiang, Manya Wadhwa, Prasann Singhal, Xinyu Zhao, Xi Ye, Kyle Mahowald, and Greg Durrett · 2024
Closest in time.
Understanding the performance gap between online and offline alignment algorithms, 2024
Yunhao Tang, Daniel Zhaohan Guo, Zeyu Zheng, Daniele Calandriello, Yuan Cao, Eugene Tarassov, Rémi Munos, Bernardo Ávila Pires, Michal Valko, Yong Cheng, and Will Dabney · 2024
Closest in time.
Efficient large language models: A survey, 2024
Zhongwei Wan, Xin Wang, Che Liu, Samiul Alam, Yu Zheng, Jiachen Liu, Zhongnan Qu, Shen Yan, Yi Zhu, Quanlu Zhang, Mosharaf Chowdhury, and Mi Zhang · 2024
Closest in time.
Quiet-star: Language models can teach themselves to think before speaking, 2024
Eric Zelikman, Georges Harik, Yijia Shao, Varuna Jayasiri, Nick Haber, and Noah D. Goodman · 2024
Closest in time.
Mmlu-cf: A contamination-free multi-task language understanding benchmark, 2024
Qihao Zhao, Yangyu Huang, Tengchao Lv, Lei Cui, Qinzheng Sun, Shaoguang Mao, Xin Zhang, Ying Xin, Qiufeng Yin, Scarlett Li, and Furu Wei · 2024
Closest in time.
Take a step back: Evoking reasoning via abstraction in large language models, 2024
Huaixiu Steven Zheng, Swaroop Mishra, Xinyun Chen, Heng-Tze Cheng, Ed H. Chi, Quoc V Le, and Denny Zhou · 2024
Closest in time.
Claude 3.7 sonnet and claude code”, 2025
Anthropic · 2025
Closest in time.
Claude 3.7 sonnet system card, 2025
Anthropic et al · 2025
Closest in time.
Training language models to reason efficiently, 2025
Daman Arora and Andrea Zanette · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI et al · 2025
Closest in time.
Ada-r1: Hybrid-cot via bi-level adaptive reasoning optimization, 2025
Haotian Luo, Haiying He, Yibo Wang, Jinluan Yang, Rui Liu, Naiqiang Tan, Xiaochun Cao, Dacheng Tao, and Li Shen · 2025
Closest in time.
Cot-valve: Length-compressible chain-of-thought tuning, 2025
Xinyin Ma, Guangnian Wan, Runpeng Yu, Gongfan Fang, and Xinchao Wang · 2025
Closest in time.
Unlocking efficient long-to-short llm reasoning with model merging, 2025
Han Wu, Yuxuan Yao, Shuqi Liu, Zehua Liu, Xiaojin Fu, Xiongwei Han, Xing Li, Hui-Ling Zhen, Tao Zhong, and Mingxuan Yuan · 2025
Closest in time.