Fetching the paper…
Reading the bibliography…
The releases of OpenAI's o-[n] series, such as o1, o3, and o4-mini, mark a significant paradigm shift in Large Language Models towards advanced reasoning capabilities.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, et al. 2020 · 1901
Earlier work this paper cites.
On the measure of intelligence
François Chollet. 2019 · 1911
Earlier work this paper cites.
The raven progressive matrices: A review of national norming studies and ethnic and socioeconomic variation within the united states
John Raven. 1989 · 1989
Earlier work this paper cites.
Towards artificial general intelligence via a multimodal foundation model
Nanyi Fei, Zhiwu Lu, Yizhao Gao, Guoxing Yang, Yuqi Huo, Jing Wen, Haoyu Lu, Ruihua Song, Xin Gao, Tao Xiang, Haoran Sun, and Jiling Wen. 2021 · 2021
Earlier work this paper cites.
True detective: A deep abductive reasoning benchmark undoable for gpt-3 and challenging for gpt-4
Maksym Del and Mark Fishel. 2023 · 2023
Earlier work this paper cites.
Unraveling the arc puzzle: Mimicking human solutions with object-centric decision transformer
Jaehyun Park, Jaegyun Im, Sanha Hwang, Mintaek Lim, Sabina Ualibekova, Sejin Kim, and Sundong Kim. 2023 · 2023
Earlier work this paper cites.
Yew Ken Chia, Vernon Toh Yan Han, Deepanway Ghosal, Lidong Bing, and Soujanya Poria. 2024 · 2024
Earlier work this paper cites.
Puzzles: A benchmark for neural algorithmic reasoning
Benjamin Estermann, Luca A. Lanzendörfer, Yannick Niedermayr, and Roger Wattenhofer. 2024 · 2024
Cited alongside, same era.
Deepanway Ghosal, Vernon Toh Yan Han, Chia Yew Ken, and Soujanya Poria. 2024 · 2024
Cited alongside, same era.
Puzzle solving using reasoning of large language models: A survey
Panagiotis Giadikiaroglou, Maria Lymperaiou, Giorgos Filandrianos, and Giorgos Stamou. 2024 · 2024
Cited alongside, same era.
Rebus: A robust evaluation benchmark of understanding symbols
Andrew Gritsevskiy, Arjun Panickssery, Aaron Kirtland, Derik Kauffman, Hans Gundlach, Irina Gritsevskaya, Joe Cavanagh, Jonathan Chiang, Lydia La Roux, and Michelle Hung. 2024 · 2024
Cited alongside, same era.
Computing power and the governance of artificial intelligence
Girish Sastry, Lennart Heim, Haydn Belfield, Markus Anderljung, Miles Brundage, Julian Hazell, Cullen O’Keefe, Gillian K. Hadfield, Richard Ngo, Konstantin Pilz, George Gor, Emma Bluemke, Sarah Shoker, Janet Egan, Robert F. Trager, Shahar Avin, Adrian Weller, Yoshua Bengio, and Diane Coyle. 2024 · 2024
Later among the works it cites.
Bongard in wonderland: Visual puzzles that still make ai go mad?
Antonia Wüst, Tim Tobiasch, Lukas Helff, Devendra S. Dhami, Constantin A. Rothkopf, and Kristian Kersting. 2024 · 2024
Later among the works it cites.
What is meant by agi? on the definition of artificial general intelligence
Bowen Xu. 2024 · 2024
Later among the works it cites.
Phd knowledge not required: A reasoning challenge for large language models
Carolyn Jane Anderson, Joydeep Biswas, Aleksander Boruch-Gruszecki, Federico Cassano, Molly Q Feldman, Arjun Guha, Francesca Lucchetti, and Zixuan Wu. 2025 · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robert Johansson. 2024 · 2024
Cited alongside, same era.
Agi: Artificial general intelligence for education
Ehsan Latif, Gengchen Mai, Matthew Nyaaba, Xuansheng Wu, Ninghao Liu, Guoyu Lu, Sheng Li, Tianming Liu, and Xiaoming Zhai. 2024 · 2024
Cited alongside, same era.
Alhassan Mumuni and Fuseini Mumuni. 2025 · 2025
Closest in time.
Enigmaeval: A benchmark of long multimodal reasoning challenges
Clinton J. Wang, Dean Lee, Cristina Menghini, Johannes Mols, Jack Doughty, Adam Khoja, Jayson Lynch, Sean Hendryx, Summer Yue, and Dan Hendrycks. 2025 · 2025
Closest in time.