Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated impressive capabilities in reasoning tasks, yet their reliance on static prompt structures and limited adaptability to complex scenarios remains a significant challenge.
On the failure to eliminate hypotheses in a conceptual task
Peter C. Wason. 1960 · 1960
Earlier work this paper cites.
A formal theory of inductive inference. part i
Ray J Solomonoff. 1964 · 1964
Earlier work this paper cites.
Judgment under uncertainty: Heuristics and biases: Biases in judgments reveal some heuristics of thinking under uncertainty
Amos Tversky and Daniel Kahneman. 1974 · 1974
Earlier work this paper cites.
Mental models: Towards a cognitive science of language, inference, and consciousness
Philip Nicholas Johnson-Laird. 1983 · 1983
Earlier work this paper cites.
The next decade in ai: four steps towards robust artificial intelligence
Gary Marcus. 2020 · 2002
Earlier work this paper cites.
Bayesian models of cognition
Thomas L Griffiths, Charles Kemp, and Joshua B Tenenbaum. 2008 · 2008
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020 · 2009
Earlier work this paper cites.
Causal models: How people think about the world and its alternatives
S Sloman. 2009 · 2009
Earlier work this paper cites.
How to grow a mind: Statistics, structure, and abstraction
Joshua B Tenenbaum, Charles Kemp, Thomas L Griffiths, and Noah D Goodman. 2011 · 2011
Earlier work this paper cites.
Complex problem solving: A case for complex cognition?
Joachim Funke. 2013 · 2013
Earlier work this paper cites.
Computational rationality: A converging paradigm for intelligence in brains, minds, and machines
Samuel J Gershman, Eric J Horvitz, and Joshua B Tenenbaum. 2015 · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022 · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. 2023 · 2023
Puzzle solving using reasoning of large language models: A survey
Panagiotis Giadikiaroglou, Maria Lymperaiou, Giorgos Filandrianos, and Giorgos Stamou. 2024 · 2024
Closest in time.
Self-explore: Enhancing mathematical reasoning in language models with fine-grained rewards
Hyeonbin Hwang, Doyoung Kim, Seungone Kim, Seonghyeon Ye, and Minjoon Seo. 2024 · 2024
Closest in time.
Learning from teaching regularization: Generalizable correlations should be easy to imitate
Can Jin, Tong Che, Hongwu Peng, Yiyuan Li, Dimitris Metaxas, and Marco Pavone. 2024 · 2024
Closest in time.
Chain of thought empowers transformers to solve inherently serial problems
Zhiyuan Li, Hong Liu, Denny Zhou, and Tengyu Ma. 2024 · 2024
Closest in time.
When is inductive inference possible?
Zhou Lu. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mathematical language models: A survey
Wentao Liu, Hanglei Hu, Jie Zhou, Yuyang Ding, Junsong Li, Jiayi Zeng, Mengliang He, Qin Chen, Bo Jiang, Aimin Zhou, et al. 2023 · 2023
Cited alongside, same era.
OpenAI. 2023 · 2023
Cited alongside, same era.
Mr-gsm8k: A meta-reasoning benchmark for large language model evaluation
Zhongshen Zeng, Pengguang Chen, Shu Liu, Haiyun Jiang, and Jiaya Jia. 2023 · 2023
Cited alongside, same era.
Medec: A benchmark for medical error detection and correction in clinical notes
Asma Ben Abacha, Wen-wai Yim, Yujuan Fu, Zhaoyi Sun, Meliha Yetisgen, Fei Xia, and Thomas Lin. 2024 · 2024
Cited alongside, same era.
Graph of thoughts: Solving elaborate problems with large language models
Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, et al. 2024 · 2024
Cited alongside, same era.
T 2 of thoughts: Temperature tree elicits reasoning in large language models
Chengkun Cai, Xu Zhao, Yucheng Du, Haoliang Liu, and Lei Li. 2024 · 2024
Cited alongside, same era.
Closest in time.
Marianna Nezhurina, Lucia Cipolina-Kun, Mehdi Cherti, and Jenia Jitsev. 2024 · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2024 · 2024
Closest in time.
Larger and more instructable language models become less reliable
Lexin Zhou, Wout Schellaert, Fernando Martínez-Plumed, Yael Moros-Daval, César Ferri, and José Hernández-Orallo. 2024 · 2024
Closest in time.
Competitive programming with large reasoning models
Ahmed El-Kishky, Alexander Wei, Andre Saraiva, Borys Minaev, Daniel Selsam, David Dohan, Francis Song, Hunter Lightman, Ignasi Clavera, Jakub Pachocki, et al. 2025 · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. 2025 · 2025
Closest in time.
Two heads are better than one: Test-time scaling of multi-agent collaborative reasoning
Can Jin, Hongwu Peng, Qixin Zhang, Yujin Tang, Dimitris N Metaxas, and Tong Che. 2025 · 2025
Closest in time.