Fetching the paper…
Reading the bibliography…
The Game Reasoning Arena library provides a framework for evaluating the decision making abilities of large language models (LLMs) through strategic board games implemented in Google OpenSpiel library.
Gamebench: Evaluating strategic reasoning abilities of llm agents, 2024
Anthony Costarelli, Mat Allen, Roman Hauksson, Grace Sodunke, Suhas Hariharan, Carlson Cheng, Wenjie Li, Joshua Clymer, and Arjun Yadav · 2024
Earlier work this paper cites.
Oguzhan Topsakal, Colby Jacob Edell, and Jackson Bailey Harper · 2024
Earlier work this paper cites.
Board Game Bench: What is board game bench? and how it works, 2025
Board Game Bench authors · 2025
Cited alongside, same era.
TextArena: Competitive text‑based games for evaluating agentic behavior in llms
Leon Guertler, Bobby Cheng, Simon Yu, Bo Liu, Leshem Choshen, and Cheston Tan · 2025
Cited alongside, same era.
lmgame-bench: How good are llms at playing games?, 2025a
Lanxiang Hu, Mingjia Huo, Yuxuan Zhang, Haoyang Yu, Eric P. Xing, Ion Stoica, Tajana Rosing, Haojian Jin, and Hao Zhang
Cited in the paper.
GameArena: Evaluating llm reasoning through live computer games
Lanxiang Hu, Qiyu Li, Anze Xie, Nan Jiang, Ion Stoica, Haojian Jin, and Hao Zhang · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…