Fetching the paper…
Reading the bibliography…
Can a Large Language Model (LLM) solve simple abstract reasoning problems? We explore this broad question through a systematic analysis of GPT on the Abstraction and Reasoning Corpus (ARC), a representative benchmark of abstract reasoning ability from limited examples in which solutions require some "core knowledge" of concepts such as objects, goal states, counting, and basic geometry.
Core knowledge
Elizabeth S Spelke and Katherine D Kinzler · 2007
Earlier work this paper cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Earlier work this paper cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2018
Earlier work this paper cites.
On the measure of intelligence
François Chollet · 2019
Earlier work this paper cites.
Dreamcoder: Growing generalizable, interpretable knowledge with wake-sleep bayesian program learning
Kevin Ellis, Catherine Wong, Maxwell Nye, Mathias Sable-Meyer, Luc Cary, Lucas Morales, Luke Hewitt, Armando Solar-Lezama, and Joshua B Tenenbaum · 2020
Earlier work this paper cites.
Solving abstract reasoning tasks with grammatical evolution
Raphael Fischer, Matthias Jakobs, Sascha Mücke, and Katharina Morik · 2020
Earlier work this paper cites.
Victor Kolev, Bogdan Georgiev, and Svetlin Penkov · 2020
Earlier work this paper cites.
Neural-guided, bidirectional program search for abstraction and reasoning
Simon Alford, Anshula Gandhi, Akshay Rangamani, Andrzej Banburski, Tony Wang, Sylee Dandekar, John Chin, Tomaso Poggio, and Peter Chin · 2021
Earlier work this paper cites.
Sébastien Ferré · 2021
Earlier work this paper cites.
Communicating natural programs to humans and machines
Sam Acquaviva, Yewen Pu, Marta Kryven, Theodoros Sechopoulos, Catherine Wong, Gabrielle Ecanow, Maxwell Nye, Michael Tessler, and Josh Tenenbaum · 2022
Earlier work this paper cites.
Object-centric compositional imagination for visual abstract reasoning
Rim Assouel, Pau Rodriguez, Perouz Taslakian, David Vazquez, and Yoshua Bengio · 2022
Earlier work this paper cites.
Arc_kaggle
Alejandro de Miquel Bleier · 2022
Earlier work this paper cites.
Pal: Program-aided language models
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig · 2022
Cited alongside, same era.
Arc-kaggle-3rd-place
Vlad Golubev · 2022
Cited alongside, same era.
Towards reasoning in large language models: A survey
Jie Huang and Kevin Chen-Chuan Chang · 2022
Cited alongside, same era.
Arc-kaggle-main
Kaggle · 2022
Cited alongside, same era.
Playgrounds for abstraction and reasoning
Subin Kim, Prin Phunyaphibarn, Donghyun Ahn, and Sundong Kim · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed H Chi, Quoc V Le, Denny Zhou, et al · 2022
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Closest in time.
Abstract visual reasoning enabled by language
Giacomo Camposampiero, Loïc Houmard, Benjamin Estermann, Joël Mathys, and Roger Wattenhofer · 2023
Closest in time.
Augmented language models: a survey
Grégoire Mialon, Roberto Dessì, Maria Lomeli, Christoforos Nalmpantis, Ram Pasunuru, Roberta Raileanu, Baptiste Rozière, Timo Schick, Jane Dwivedi-Yu, Asli Celikyilmaz, et al · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Arc-kaggle-5th-place
Agnis Liukis · 2022
Cited alongside, same era.
Evaluating understanding on conceptual abstraction benchmarks
Victor Vikram Odouard and Melanie Mitchell · 2022
Cited alongside, same era.
Arc-kaggle-8th-place
Andy Penrose · 2022
Cited alongside, same era.
Reasoning with language model prompting: A survey
Shuofei Qiao, Yixin Ou, Ningyu Zhang, Xiang Chen, Yunzhi Yao, Shumin Deng, Chuanqi Tan, Fei Huang, and Huajun Chen · 2022
Cited alongside, same era.
Arc-solution
top quarks · 2022
Cited alongside, same era.
OpenAI
Cited in the paper.
Suvir Mirchandani, Fei Xia, Pete Florence, Brian Ichter, Danny Driess, Montserrat Gonzalez Arenas, Kanishka Rao, Dorsa Sadigh, and Andy Zeng · 2023
Closest in time.
The conceptarc benchmark: Evaluating understanding and generalization in the arc domain
Arseny Moskvichev, Victor Vikram Odouard, and Melanie Mitchell · 2023
Closest in time.
Larc solving with gpt4
Yewen Pu and Simon Alford · 2023
Closest in time.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
Hypothesis search: Inductive reasoning with language models
Ruocheng Wang, Eric Zelikman, Gabriel Poesia, Yewen Pu, Nick Haber, and Noah D Goodman · 2023
Closest in time.
Graphs, constraints, and search for the abstraction and reasoning corpus
Yudong Xu, Elias B. Khalil, and Scott Sanner · 2023
Closest in time.