Fetching the paper…
Reading the bibliography…
Existing reasoning benchmarks for large language models (LLMs) frequently fail to capture authentic creativity, often rewarding memorization of previously observed patterns.
On the measure of intelligence, 2019
F. Chollet · 1911
Earlier work this paper cites.
Difficulty rating of sudoku puzzles by a computational model
R. Pelánek · 2011
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
R. Lowe, Y. Wu, A. Tamar, J. Harb, P. Abbeel, and I. Mordatch · 2017
Earlier work this paper cites.
Deep q-learning from demonstrations
T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, D. Horgan, J. Quan, A. Sendonaris, I. Osband, G. Dulac-Arnold, J. Agapiou, J. Z. Leibo, and A. Gruslys · 2018
Earlier work this paper cites.
Recurrent relational networks
R. Palm, U. Paquet, and O. Winther · 2018
Earlier work this paper cites.
Satnet: Bridging deep learning and logical reasoning using a differentiable satisfiability solver
P.-W. Wang, P. Donti, B. Wilder, and Z. Kolter · 2019
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Sudokupad, 2021
S. Neumann · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V. Le, and D. Zhou · 2022
Earlier work this paper cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. A. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y.-F. Li, S. M. Lundberg, H. Nori, H. Palangi, M. T. Ribeiro, and Y. Zhang · 2023
Cited alongside, same era.
Large language model guided tree-of-thought
J. Long · 2023
Cited alongside, same era.
Puzzles: A benchmark for neural algorithmic reasoning
B. Estermann, L. A. Lanzendörfer, Y. Niedermayr, and R. Wattenhofer · 2024
Cited alongside, same era.
Puzzle solving using reasoning of large language models: A survey
P. Giadikiaroglou, M. Lymperaiou, G. Filandrianos, and G. Stamou · 2024
Cited alongside, same era.
Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai, 2024
Beyond autoregression: Discrete diffusion for complex reasoning and planning
J. Ye, J. Gao, S. Gong, L. Zheng, X. Jiang, Z. Li, and L. Kong · 2024
Later among the works it cites.
A careful examination of large language model performance on grade school arithmetic
H. Zhang, J. Da, D. Lee, V. Robinson, C. Wu, W. Song, T. Zhao, P. Raja, C. Zhuang, D. Slack, et al · 2024
Later among the works it cites.
https://logic-masters.de
Logic masters germany · 2025
Closest in time.
Train for the worst, plan for the best: Understanding token ordering in masked diffusions
J. Kim, K. Shah, V. Kontonis, S. Kakade, and S. Chen · 2025
Closest in time.
Llms can easily learn to reason from demonstrations structure, not content, is what matters!
D. Li, S. Cao, T. Griggs, S. Liu, X. Mo, E. Tang, S. Hegde, K. Hakhamaneshi, S. G. Patil, M. Zaharia, et al · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Glazer, E. Erdil, T. Besiroglu, D. Chicharro, E. Chen, A. Gunning, C. F. Olsson, J.-S. Denain, A. Ho, E. de Oliveira Santos, O. Järviniemi, M. Barnett, R. Sandler, M. Vrzala, J. Sevilla, Q. Ren, E. Pratt, L. Levine, G. Barkley, N. Stewart, B. Grechuk, T. Grechuk, S. V. Enugandla, and M. Wildon · 2024
Cited alongside, same era.
Logicgame: Benchmarking rule-based reasoning abilities of large language models
J. Gui, Y. Liu, J. Cheng, X. Gu, X. Liu, H. Wang, Y. Dong, J. Tang, and M. Huang · 2024
Cited alongside, same era.
PuzzlePlex: A benchmark to evaluate the reasoning and planning of large language models on puzzles
Y. Long, T. Jiang, Y. Zhao, A. Cohan, and D. Shasha · 2024
Cited alongside, same era.
Artificial kuramoto oscillatory neurons
T. Miyato, S. Löwe, A. Geiger, and M. Welling · 2024
Cited alongside, same era.
Balrog: Benchmarking agentic llm and vlm reasoning on games
D. Paglieri, B. Cupiał, S. Coward, U. Piterbarg, M. Wolczyk, A. Khan, E. Pignatelli, Ł. Kuciński, L. Pinto, R. Fergus, et al · 2024
Cited alongside, same era.
Causal language modeling can elicit search and reasoning capabilities on logic puzzles
K. Shah, N. Dikkala, X. Wang, and R. Panigrahy · 2024
Cited alongside, same era.
Step-by-step reasoning to solve grid puzzles: Where do LLMs falter?
N. Tyagi, M. Parmar, M. Kulkarni, A. Rrv, N. Patel, M. Nakamura, A. Mitra, and C. Baral · 2024
Cited alongside, same era.
Closest in time.
N. Muennighoff, Z. Yang, W. Shi, X. L. Li, L. Fei-Fei, H. Hajishirzi, L. Zettlemoyer, P. Liang, E. Candès, and T. Hashimoto · 2025
Closest in time.
OpenAI o3 and o4-mini System Card
OpenAI · 2025
Closest in time.
L. Phan, A. Gatti, Z. Han, N. Li, J. Hu, H. Zhang, C. B. C. Zhang, M. Shaaban, J. Ling, S. Shi, et al · 2025
Closest in time.
Vgrp-bench: Visual grid reasoning puzzle benchmark for large vision-language models
Y. Ren, K. Tertikas, S. Maiti, J. Han, T. Zhang, S. Süsstrunk, and F. Kokkinos · 2025
Closest in time.
Enigmaeval: A benchmark of long multimodal reasoning challenges, 2025
C. J. Wang, D. Lee, C. Menghini, J. Mols, J. Doughty, A. Khoja, J. Lynch, S. Hendryx, S. Yue, and D. Hendrycks · 2025
Closest in time.