Fetching the paper…
Reading the bibliography…
Playing video games requires perception, memory, and planning, exactly the faculties modern large language model (LLM) agents are expected to master.
The probable error of a mean
Gosset, W.S.: · 1908
Earlier work this paper cites.
A new measure of rank correlation
Kendall, M.G.: · 1938
Earlier work this paper cites.
Primary, secondary, and meta-analysis of research
Glass, G.V.: · 1976
Earlier work this paper cites.
Temporal difference learning and td-gammon
Tesauro, G., et al.: · 1995
Earlier work this paper cites.
Active learning literature survey
Settles, B.: · 1995
Earlier work this paper cites.
Sokoban is pspace-complete
Culberson, J.: · 1997
Earlier work this paper cites.
SOKOBAN and other motion planning problems
Dor, D., Zwick, U.: · 1999
Earlier work this paper cites.
Human-level ai’s killer application: Interactive computer games
Laird, J.E., van Lent, M.: · 2001
Earlier work this paper cites.
Tetris is hard, even to approximate
Demaine, E.D., Hohenberger, S., Liben - · 2003
Earlier work this paper cites.
Complexity of planning with partial observability
Rintanen, J.: · 2004
Earlier work this paper cites.
A similarity measure for indefinite rankings
Webber, W., Moffat, A., Zobel, J.: · 2010
Earlier work this paper cites.
Minimax and expectimax algorithm to solve 2048
Zaky, A.: · 2014
Earlier work this paper cites.
Bejeweled, candy crush and other match -
Gualà, S., Leucci, S., Natale, E.: · 2014
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., Zaremba, W.: · 2016
Earlier work this paper cites.
Selective association between tetris game play and visuospatial working memory: A preliminary investigation
Lau-Zhu, A., Holmes, E.A., Butterfield, S., Holmes, J.: · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al.: · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., Klimov, O.: · 2017
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Carlini, N., Liu, C., Erlingsson, Ú., Kos, J., Song, D.: · 2018
Earlier work this paper cites.
gym-sokoban
Schrader, M.P.B.: · 2018
Earlier work this paper cites.
General video game ai: A multi-track framework for evaluating agents, games and content generation algorithms
Perez-Liebana, D., Liu, J., Khalifa, A., Gaina, R.D., Togelius, J., Lucas, S.M.: · 2019
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Kadavath, S., Arora, P., Basart, S., Tang, D.S., et al.: · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., Schulman, J.: · 2021
Earlier work this paper cites.
Webshop: Towards scalable real-world web interaction with grounded language agents
Yao, S., Chen, H., Yang, J., Narasimhan, K.: · 2022
Earlier work this paper cites.
Smartplay: A benchmark for llms as intelligent agents
Wu, Y., Tang, X., Mitchell, T.M., Li, Y.: · 2023
Earlier work this paper cites.
Gpqa: Graded physics question answering benchmark for large language models
Rein, D., Hou, B.L., Stickland, A.C., Petty, J., Pang, R.Y., Dirani, J., Michael, J., Bowman, S.R.: · 2023
Earlier work this paper cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y., et al.: · 2023
Earlier work this paper cites.
Tetrisrl: Reinforcement learning for tetris
TetrisRL: · 2023
Earlier work this paper cites.
Gameeval: Evaluating llms on conversational games
Qiao, D., Wu, C., Liang, Y., Li, J., Duan, N.: · 2023
Earlier work this paper cites.
Swe-bench: Can language models resolve real-world github issues?
Jimenez, C.E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., Narasimhan, K.: · 2023
Earlier work this paper cites.
Webarena: A realistic web environment for building autonomous agents
Zhou, S., Xu, F.F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., et al.: · 2023
Earlier work this paper cites.
Gaia: a benchmark for general ai assistants
Mialon, G., Fourrier, C., Wolf, T., LeCun, Y., Scialom, T.: · 2023
Cited alongside, same era.
Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change
Valmeekam, K., Marquez, M., Olmo, A., Sreedharan, S., Kambhampati, S.: · 2023
Cited alongside, same era.
Gymnasium: A standard interface for reinforcement learning environments
Towers, M., Kwiatkowski, A., Terry, J., Balis, J.U., De Cola, G., Deleu, T., Goulão, M., Kallinteris, A., Krimmel, M., KG, A., et al.: · 2024
Cited alongside, same era.
Fine-tuning large vision-language models as decision-making agents via reinforcement learning
Zhai, Y., Bai, H., Lin, Z., Pan, J., Tong, S., Zhou, Y., Suhr, A., Xie, S., LeCun, Y., Ma, Y., Levine, S.: · 2024
Cited alongside, same era.
Balrog: Benchmarking agentic llm and vlm reasoning on games
Paglieri, D., Cupiał, B., Coward, S., Piterbarg, U., Wolczyk, M., Khan, A., Pignatelli, E., Kuciński, Ł., Pinto, L., Fergus, R., et al.: · 2024
Sweet-rl: Training multi-turn llm agents on collaborative reasoning tasks
Zhou, Y., Jiang, S., Tian, Y., Weston, J., Levine, S., Sukhbaatar, S., Li, X.: · 2025
Closest in time.
Korgym: A dynamic game platform for llm reasoning evaluation (2025)
Shi, J., Yang, J., Liu, J., Bu, X., Chen, J., Zhou, J., Ma, K., Wen, Z., Wang, B., He, Y., Song, L., Zhu, H., Li, S., Wang, X., Zhang, W., Yuan, R., Yao, Y., Yang, W., Wang, Y., Fang, S., Yuan, S., He, Q., Tang, X., Tan, Y., Zhou, W., Zhang, Z., Li, Z., Huang, W., Zhang, G.: · 2025
Closest in time.
Lmact: A benchmark for in-context imitation learning with long multimodal demonstrations (2025)
Ruoss, A., Pardo, F., Chan, H., Li, B., Mnih, V., Genewein, T.: · 2025
Closest in time.
Enigmaeval: A benchmark of long multimodal reasoning challenges
Wang, C.J., Lee, D., Menghini, C., Mols, J., Doughty, J., Khoja, A., Lynch, J., Hendryx, S., Yue, S., Hendrycks, D.: · 2025
Closest in time.
Claude 3.7 sonnet: Frontier reasoning made practical (February 2025) Accessed: 2025-05-02
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gamebench: Evaluating strategic reasoning abilities of llm agents
Costarelli, A., Allen, M., Hauksson, R., Sodunke, G., Hariharan, S., Cheng, C., Li, W., Clymer, J., Yadav, A.: · 2024
Cited alongside, same era.
Thinking in space: How multimodal large language models see, remember, and recall spaces
Yang, J., Yang, S., Gupta, A.W., Han, R., Fei-Fei, L., Xie, S.: · 2024
Cited alongside, same era.
Waytowich, N.R., White, D., Sunbeam, M., Goecks, V.G.: · 2024
Cited alongside, same era.
Mosquera, M., Pinzon, J.S., Rios, M., Fonseca, Y., Giraldo, L.F., Quijano, N., Manrique, R.: · 2024
Cited alongside, same era.
Openai o1 system card (December 2024) arXiv preprint arXiv:2412.16720
OpenAI: · 2024
Cited alongside, same era.
Ufo: A unified framework for gui interaction in windows applications
Group, A.R.: · 2024
Cited alongside, same era.
Dong, Y., Jiang, X., Liu, H., Jin, Z., Gu, B., Yang, M., Li, G.: · 2024
Cited alongside, same era.
Anthropic: · 2025
Closest in time.
Openai o3 and o4-mini system card (April 2025) Accessed: 2025-05-10
OpenAI: · 2025
Closest in time.
Gemini 2.5: Our most intelligent ai model (March 2025) Accessed: 2025-05-10
DeepMind, G.: · 2025
Closest in time.
Infiguiagent: A multimodal agent for gui interaction and reasoning
Group, A.R.: · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al.: · 2025
Closest in time.
Mmlu-pro benchmark leaderboard
VALS AI: · 2025
Closest in time.
Humanity’s last exam leaderboard
Scale AI: · 2025
Closest in time.
Humanity’s last exam leaderboard (text only)
Scale AI: · 2025
Closest in time.
Gpqa benchmark leaderboard
VALS AI: · 2025
Closest in time.
Phybench: Holistic evaluation of physical perception and reasoning in large language models
Qiu, S., Guo, S., Zhuo, Y., Wang, Y., Li, Z., Zhang, Y., Wang, Y., Li, Z., Zhang, Y., Wang, Y., Li, Z., Zhang, Y.: · 2025
Closest in time.
Can mllms reason in multimodality? emma: An enhanced multimodal reasoning benchmark
Hao, Y., Gu, J., Wang, H.W., Li, L., Yang, Z., Wang, L., Cheng, Y.: · 2025
Closest in time.
Math 500 benchmark leaderboard
Vals AI: · 2025
Closest in time.
Aime benchmark leaderboard
Vals AI: · 2025
Closest in time.
Livebench leaderboard
LiveBench Team: · 2025
Closest in time.
Aider llm leaderboards
Aider Team: · 2025
Closest in time.
Bigcodebench leaderboard
BigCodeBench Team: · 2025
Closest in time.
Vista: Visual language understanding benchmark leaderboard
Scale AI: · 2025
Closest in time.
Chatbot arena leaderboard
LMSYS Org: · 2025
Closest in time.
Mmmu benchmark leaderboard
Vals AI: · 2025
Closest in time.
Sirdeshmukh, V., Deshpande, K., Mols, J., Jin, L., Cardona, E.Y., Lee, D., Kritz, J., Primack, W., Yue, S., Xing, C.: · 2025
Closest in time.
Multichallenge leaderboard
Scale AI: · 2025
Closest in time.
Enigmaeval benchmark leaderboard
Scale AI: · 2025
Closest in time.
Nyt connections benchmark: Evaluating llms with extended word association puzzles
Mazur, L.: · 2025
Closest in time.
Hopkins, J., Bakler, M., Khan, A.: · 2025
Closest in time.
Textgames: Learning to self-play text-based puzzle games via language model reasoning
Hudi, F., Winata, G.I., Zhang, R., Aji, A.F.: · 2025
Closest in time.
Dapo: An open-source llm reinforcement learning system at scale
Yu, Q., Zhang, Z., Zhu, R., Yuan, Y., Zuo, X., Yue, Y., Fan, T., Liu, G., Liu, L., Liu, X., et al.: · 2025
Closest in time.