Fetching the paper…
Reading the bibliography…
We present gg-bench, a collection of game environments designed to evaluate general reasoning capabilities in language models.
Interactive fiction games: A colossal adventure
Matthew Hausknecht, Prithviraj Ammanabrolu, Côté Marc-Alexandre, and Yuan Xingdi · 1909
Earlier work this paper cites.
On the measure of intelligence, 2019
François Chollet · 1911
Earlier work this paper cites.
A proposal for the Dartmouth summer research project on artificial intelligence
John McCarthy, Marvin L Minsky, Nathaniel Rochester, and Claude E Shannon · 1955
Earlier work this paper cites.
Heuristic DENDRAL: A program for generating explanatory hypotheses
Bruce Buchanan, Georgia Sutherland, and Edward A Feigenbaum · 1969
Earlier work this paper cites.
What is intelligence?: Contemporary viewpoints on its nature and definition
Robert J Sternberg and Douglas K Detterman · 1986
Earlier work this paper cites.
Cyc: toward programs with common sense
Douglas B Lenat, Ramanathan V. Guha, Karen Pittman, Dexter Pratt, and Mary Shepherd · 1990
Earlier work this paper cites.
Deep Blue
Murray Campbell, A.Joseph Hoane, and Feng hsiung Hsu · 2002
Earlier work this paper cites.
Winnowing: local algorithms for document fingerprinting
Saul Schleimer, Daniel S Wilkerson, and Alex Aiken · 2003
Earlier work this paper cites.
Artificial general intelligence , volume 2
Ben Goertzel and Cassio Pennachin · 2007
Earlier work this paper cites.
A collection of definitions of intelligence
Shane Legg, Marcus Hutter, et al · 2007
Earlier work this paper cites.
Frames of mind: The theory of multiple intelligences
Howard E Gardner · 2011
Earlier work this paper cites.
Measuring intelligence through games, 2011
Tom Schaul, Julian Togelius, and Jürgen Schmidhuber · 2011
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Deep reinforcement learning from self-play in imperfect-information games, 2016
Johannes Heinrich and David Silver · 2016
Earlier work this paper cites.
Artificial intelligence: a modern approach
Stuart J Russell and Peter Norvig · 2016
Cited alongside, same era.
Proximal policy optimization algorithms, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm, 2017
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2017
Cited alongside, same era.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
Noam Brown and Tuomas Sandholm · 2018
Cited alongside, same era.
Dynabench: Rethinking benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams · 2021
ReAct: Synergizing reasoning and acting in language models, 2023
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao · 2023
Later among the works it cites.
Chatbot arena: An open platform for evaluating LLMs by human preference
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica · 2024
Later among the works it cites.
GameBench: Evaluating strategic reasoning abilities of LLM agents, 2024
Anthony Costarelli, Mat Allen, Roman Hauksson, Grace Sodunke, Suhas Hariharan, Carlson Cheng, Wenjie Li, Joshua Clymer, and Arjun Yadav · 2024
Later among the works it cites.
SWE-bench: Can language models resolve real-world GitHub issues?, 2024
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al · 2021
Cited alongside, same era.
Stable-Baselines3: Reliable reinforcement learning implementations
Antonin Raffin, Ashley Hill, Adam Gleave, Anssi Kanervisto, Maximilian Ernestus, and Noah Dormann · 2021
Cited alongside, same era.
Neural theory-of-mind? on the limits of social intelligence in large LMs
Maarten Sap, Ronan Le Bras, Daniel Fried, and Yejin Choi · 2022
Cited alongside, same era.
Challenging BIG-Bench tasks and whether chain-of-thought can solve them, 2022
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny Zhou, and Jason Wei · 2022
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with GPT-4, 2023
Sébastien Bubeck, Varun Chadrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Cited alongside, same era.
PAL: Program-aided language models
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig · 2023
Cited alongside, same era.
Reflexion: Language agents with verbal reinforcement learning, 2023
Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2023
Cited alongside, same era.
Discovering and exploring cases of educational source code plagiarism with Dolos
Rien Maertens, Maarten Van Neyghem, Maxiem Geldhof, Charlotte Van Petegem, Niko Strijbol, Peter Dawyndt, and Bart Mesuere · 2024
Later among the works it cites.
Beyond accuracy: Evaluating self-consistency of code large language models with IdentityChain
Marcus J. Min, Yangruibo Ding, Luca Buratti, Saurabh Pujar, Gail Kaiser, Suman Jana, and Baishakhi Ray · 2024
Later among the works it cites.
Introducing OpenAI o1, 2024
OpenAI · 2024
Later among the works it cites.
Oguzhan Topsakal, Colby Jacob Edell, and Jackson Bailey Harper · 2024
Later among the works it cites.
τ \tau -bench: A benchmark for tool-agent-user interaction in real-world domains, 2024
Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan · 2024
Later among the works it cites.
ZeroSumEval: An extensible framework for scaling llm evaluation with inter-model competition
Hisham A Alyahya, Haidar Khan, Yazeed Alnumay, M Saiful Bari, and Bülent Yener · 2025
Closest in time.
Gemini 2.5 Pro
Google DeepMind · 2025
Closest in time.
VideoGameBench: Research preview
Alex Zhang and Ofir Press · 2025
Closest in time.
Absolute Zero: Reinforced self-play reasoning with zero data
Andrew Zhao, Yiran Wu, Yang Yue, Tong Wu, Quentin Xu, Matthieu Lin, Shenzhi Wang, Qingyun Wu, Zilong Zheng, and Gao Huang · 2025
Closest in time.