Fetching the paper…
Reading the bibliography…
We explore the abstract reasoning abilities of text-only and multimodal versions of GPT-4, using the ConceptARC benchmark [10], which is designed to evaluate robust understanding and reasoning with core-knowledge concepts.
Core knowledge
E. S. Spelke and K. D. Kinzler · 2007
Earlier work this paper cites.
Toddlers infer higher-order relational principles in causal learning
C. M. Walker and A. Gopnik · 2014
Earlier work this paper cites.
On the measure of intelligence
F. Chollet · 2019
Earlier work this paper cites.
Fast and flexible: Human program induction in abstract reasoning tasks
A. Johnson, W. K. Vong, B. M. Lake, and T. M. Gureckis · 2021
Earlier work this paper cites.
Impact of pretraining term frequencies on few-shot numerical reasoning
Y. Razeghi, R. L. Logan IV, M. Gardner, and S. Singh · 2022
Earlier work this paper cites.
Emergent abilities of large language models
J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus · 2022
Earlier work this paper cites.
The Abstraction and Reasoning Corpus (ARC)
F. Chollet · 2023
Earlier work this paper cites.
Finishing 2nd in Kaggle’s Abstraction and Reasoning Challenge
A. de Miquel Bleier · 2023
Cited alongside, same era.
Large language models are not strong abstract reasoners
G. Gendron, Q. Bao, M. Witbrock, and G. Dobbie · 2023
Cited alongside, same era.
Kaggle Abstraction and Reasoning Challenge
Kaggle.com · 2023
Cited alongside, same era.
Can llms really reason and plan?
S. Kambhampati · 2023
Cited alongside, same era.
R. T. McCoy, S. Yao, D. Friedman, M. Hardy, and T. L. Griffiths · 2023
Cited alongside, same era.
Large language models as general pattern machines
The conceptarc benchmark: Evaluating understanding and generalization in the arc domain
A. Moskvichev, V. V. Odouard, and M. Mitchell · 2023
Closest in time.
Hypothesis search: Inductive reasoning with language models
R. Wang, E. Zelikman, G. Poesia, Y. Pu, N. Haber, and N. D. Goodman · 2023
Closest in time.
Emergent analogical reasoning in large language models
T. Webb, K. J. Holyoak, and H. Lu · 2023
Closest in time.
1st place solution + code and official documentation
J. S. Wind · 2023
Closest in time.
Z. Wu, L. Qiu, A. Ross, E. Akyürek, B. Chen, B. Wang, N. Kim, J. Andreas, and Y. Kim · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Mirchandani, F. Xia, P. Florence, B. Ichter, D. Driess, M. G. Arenas, K. Rao, D. Sadigh, and A. Zeng · 2023
Cited alongside, same era.
Y. Xu, W. Li, P. Vaezipoor, S. Sanner, and E. B. Khalil · 2023
Closest in time.