Fetching the paper…
Reading the bibliography…
We introduce a novel and extensible benchmark for large language models (LLMs) through grid-based games such as Tic-Tac-Toe, Connect Four, and Gomoku.
Hellaswag: Can a machine really finish your sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 1905
Earlier work this paper cites.
Go-moku solved by new search techniques
L. V. Allis, H. J. van den Herik, and M. P. Huntjens · 1996
Earlier work this paper cites.
Artificial General Intelligence
B. Goertzel and C. Pennachin, editors · 2007
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
S. Lin, J. Hilton, and O. Evans · 2021
Earlier work this paper cites.
A comprehensive overview of large language models
H. Naveed, A. U. Khan, S. Qiu, M. Saqib, S. Anwar, M. Usman, N. Akhtar, N. Barnes, and A. Mian · 2023
Earlier work this paper cites.
Challenges and applications of large language models
J. Kaddour, J. Harris, M. Mozes, H. Bradley, R. Raileanu, and R. McHardy · 2023
Earlier work this paper cites.
Evaluating large language models: A comprehensive survey
Z. Guo, R. Jin, C. Liu, Y. Huang, D. Shi, L. Yu, et al · 2023
Earlier work this paper cites.
Testing spatial reasoning of large language models: The case of tic-tac-toe
D. Liga and L. Pasetto · 2023
Earlier work this paper cites.
A survey of large language models
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, et al · 2023
Earlier work this paper cites.
Nvidia ceo predicts agi in 5 years, 2024
J. Huang · 2024
Earlier work this paper cites.
Meta ai chief skeptical about agi, quantum computing, 2023
Y. LeCun · 2024
Earlier work this paper cites.
The exciting, perilous journey toward agi, 2024
I. Sutskever · 2024
Earlier work this paper cites.
Improving language understanding by generative pre-training, 2024
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2024
Earlier work this paper cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2024
Earlier work this paper cites.
Gemini: A family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, Y. Wu, J. B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2024
Cited alongside, same era.
Model card and evaluations for claude models, 2024
Anthropic · 2024
Cited alongside, same era.
Grok large language model, 2024
Grok Large Language Model · 2024
Cited alongside, same era.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2024
Cited alongside, same era.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al · 2024
Cited alongside, same era.
Smartplay: A benchmark for llms as intelligent agents
Y. Wu, X. Tang, T. M. Mitchell, and Y. Li · 2024
Closest in time.
Mindagent: Emergent gaming interaction
R. Gong, Q. Huang, X. Ma, H. Vo, Z. Durante, Y. Noda, Z. Zheng, S.-C. Zhu, D. Terzopoulos, L. Fei-Fei, et al · 2024
Closest in time.
Playing repeated games with large language models
E. Akata, L. Schulz, J. Coda-Forno, S. J. Oh, M. Bethge, and E. Schulz · 2024
Closest in time.
Strategic behavior of large language models: Game structure vs. contextual framing
N. Lorè and B. Heydari · 2024
Closest in time.
Can large language models play text games well? current state-of-the-art and open questions
C. F. Tsai, X. Zhou, S. S. Liu, J. Li, M. Yu, and H. Mei · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Large language models: A survey
S. Minaee, T. Mikolov, N. Nikzad, M. Chenaghlu, R. Socher, X. Amatriain, and J. Gao · 2024
Cited alongside, same era.
A survey on evaluation of large language models
Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang, et al · 2024
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 2024
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 2024
Cited alongside, same era.
Holistic evaluation of language models
P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Kumar, et al · 2024
Cited alongside, same era.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2024
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonso, et al · 2024
Cited alongside, same era.
Closest in time.
Can large language models serve as rational players in game theory? a systematic analysis
C. Fan, J. Chen, Y. Jin, and H. He · 2024
Closest in time.
Gtbench: Uncovering the strategic reasoning limitations of llms via game-theoretic evaluations
J. Duan, R. Zhang, J. Diffenderfer, B. Kailkhura, L. Sun, E. Stengel-Eskin, et al · 2024
Closest in time.
Game-bench: Evaluating strategic reasoning abilities of llm agents
A. Costarelli, M. Allen, R. Hauksson, G. Sodunke, S. Hariharan, C. Cheng, et al · 2024
Closest in time.
Benchmarking large language model (llm) performance for game playing via tic-tac-toe
O. Topsakal and J. B. Harper · 2024
Closest in time.
Large language models and games: A survey and roadmap
R. Gallotta, G. Todd, M. Zammit, S. Earle, A. Liapis, J. Togelius, and G. N. Yannakakis · 2024
Closest in time.
A survey on large language model-based game agents
S. Hu, T. Huang, F. Ilhan, S. Tekin, G. Liu, R. Kompella, and L. Liu · 2024
Closest in time.
Llm game benchmark, 2024
LLM Game Benchmark · 2024
Closest in time.
Available online: https://en.wikipedia.org/wiki/Tic-tac-toe (accessed on 7 June 2024)
Tic tac toe game, 2024 · 2024
Closest in time.
Available online: https://en.wikipedia.org/wiki/Connect_Four (accessed on 7 June 2024)
Connect4, 2024 · 2024
Closest in time.
Available online: https://en.wikipedia.org/wiki/Gomoku (accessed on 7 June 2024)
Gomoku, 2024 · 2024
Closest in time.
Large language models: A survey
S. Minaee, T. Mikolov, N. Nikzad, M. Chenaghlu, R. Socher, X. Amatriain, and J. Gao · 2024
Closest in time.