Fetching the paper…
Reading the bibliography…
Contemporary large language models are powerful problem-solving tools, but they exhibit weaknesses in their reasoning abilities which ongoing research seeks to mitigate.
Arkoudas, K · 2023
Earlier work this paper cites.
Measuring and narrowing the compositionality gap in language models
Press, O., Zhang, M., Min, S., Schmidt, L., Smith, N., and Lewis, M · 2023
Earlier work this paper cites.
Large language models can be easily distracted by irrelevant context
Shi, F., Chen, X., Misra, K., Scales, N., Dohan, D., Chi, E. H., Schärli, N., and Zhou, D · 2023
Earlier work this paper cites.
Claude 3.5 Sonnet, 2024
Anthropic · 2024
Earlier work this paper cites.
Faith and fate: Limits of transformers on compositionality
Dziri, N., Lu, X., Sclar, M., Li, X. L., Jiang, L., Lin, B. Y., Welleck, S., West, P., Bhagavatula, C., Le Bras, R., et al · 2024
Earlier work this paper cites.
Language models, like humans, show content effects on reasoning tasks
Lampinen, A. K., Dasgupta, I., Chan, S. C., Sheahan, H. R., Creswell, A., Kumaran, D., McClelland, J. L., and Hill, F · 2024
Earlier work this paper cites.
Introducing Llama 3.1: Our most capable models to date, 2024
Meta · 2024
Cited alongside, same era.
GSM-Symbolic: Understanding the limitations of mathematical reasoning in large language models
Mirzadeh, I., Alizadeh, K., Shahrokhi, H., Tuzel, O., Bengio, S., and Farajtabar, M · 2024
Cited alongside, same era.
A comprehensive overview of large language models
Naveed, H., Khan, A. U., Qiu, S., Saqib, M., Anwar, S., Usman, M., Akhtar, N., Barnes, N., and Mian, A · 2024
Cited alongside, same era.
Our next-generation model: Gemini 1.5, 2024
Pichai, S. and Hassabis, D · 2024
Cited alongside, same era.
Salinas, A. and Morstatter, F · 2024
Cited alongside, same era.
LLMs still can’t plan; can LRMs? a preliminary evaluation of OpenAI’s o1 on PlanBench
Valmeekam, K., Stechly, K., and Kambhampati, S · 2024
Later among the works it cites.
Grokked transformers are implicit reasoners: A mechanistic journey to the edge of generalization
Wang, B., Yue, X., Su, Y., and Sun, H · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K · 2024
Later among the works it cites.
Larger and more instructable language models become less reliable
Zhou, L., Schellaert, W., Martínez-Plumed, F., Moros-Daval, Y., Ferri, C., and Hernández-Orallo, J · 2024
Later among the works it cites.
Kimi k1.5: Scaling reinforcement learning with LLMs, 2025
Kimi Team · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the self-verification limitations of large language models on reasoning and planning tasks
Stechly, K., Valmeekam, K., and Kambhampati, S · 2024
Cited alongside, same era.
DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning, 2025a
DeepSeek-AI
Cited in the paper.
DeepSeek-R1 model card, 2025b
DeepSeek-AI
Cited in the paper.
Hello GPT-4o, 2024a
OpenAI
Cited in the paper.
Learning to reason with LLMs, 2024b
OpenAI
Cited in the paper.
OpenAI o1-mini, 2024c
OpenAI
Cited in the paper.