Fetching the paper…
Reading the bibliography…
Decision-making is a complex process requiring diverse abilities, making it an excellent framework for evaluating Large Language Models (LLMs).
Equilibrium points in n-person games
John F Nash · 1950
Earlier work this paper cites.
Non-cooperative games
John F Nash · 1951
Earlier work this paper cites.
The pure theory of public expenditure
Paul A Samuelson · 1954
Earlier work this paper cites.
Counterspeculation, auctions, and competitive sealed tenders
William Vickrey · 1961
Earlier work this paper cites.
Pure competition, coalitional power, and fair division
Lloyd S Shapley and Martin Shubik · 1969
Earlier work this paper cites.
The sequential truel
D Mark Kilgour · 1975
Earlier work this paper cites.
Equilibrium points of infinite sequential truels
D Marc Kilgour · 1977
Earlier work this paper cites.
Concours résultats complets. les victimes se sont plu à jouer le 14 d’atout
Alain Ledoux · 1981
Earlier work this paper cites.
Auctions and bidding
R Preston McAfee and John McMillan · 1987
Earlier work this paper cites.
The Ecology of Computation
Bernardo A. Huberman · 1988
Earlier work this paper cites.
Inductive reasoning and bounded rationality
W Brian Arthur · 1994
Earlier work this paper cites.
The dynamics of social dilemmas
Natalie S Glance and Bernardo A Huberman · 1994
Earlier work this paper cites.
Unraveling in guessing games: An experimental study
Rosemarie Nagel · 1995
Earlier work this paper cites.
Retrospectives: The ethology of homo economicus
Joseph Persky · 1995
Earlier work this paper cites.
The truel
D Marc Kilgour and Steven J Brams · 1997
Earlier work this paper cites.
Representations and solutions for game-theoretic problems
Daphne Koller and Avi Pfeffer · 1997
Earlier work this paper cites.
The theory of institutional design
Robert E Goodin · 1998
Earlier work this paper cites.
A puzzle for pirates
Ian Stewart · 1999
Earlier work this paper cites.
What is rational about nash equilibria?
Mathias Risse · 2000
Earlier work this paper cites.
Instinctive and cognitive reasoning: A study of response times
Ariel Rubinstein · 2007
Earlier work this paper cites.
Game theory
Roger B Myerson · 2013
Earlier work this paper cites.
Generalized divide the dollar
Daniel Ashlock and Garrison Greenwood · 2016
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Earlier work this paper cites.
Introducing chatgpt
OpenAI · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Evaluating multi-agent coordination abilities in large language models
Saaket Agashe, Yue Fan, and Xin Eric Wang · 2023
Cited alongside, same era.
Using large language models to simulate multiple humans and replicate human subject studies
Gati V Aher, Rosa I Arriaga, and Adam Tauman Kalai · 2023
Cited alongside, same era.
Playing repeated games with large language models
Elif Akata, Lion Schulz, Julian Coda-Forno, Seong Joon Oh, Matthias Bethge, and Eric Schulz · 2023
Cited alongside, same era.
Education in the era of generative artificial intelligence (ai): Understanding the potential benefits of chatgpt in promoting teaching and learning
David Baidoo-Anu and Leticia Owusu Ansah · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Chatgpt or grammarly? evaluating chatgpt on grammatical error correction benchmark
Haoran Wu, Wenxuan Wang, Yuxuan Wan, Wenxiang Jiao, and Michael Lyu · 2023
Later among the works it cites.
Exploring large language models for communication games: An empirical study on werewolf
Yuzhuang Xu, Shuo Wang, Peng Li, Fuwen Luo, Xiaolong Wang, Weidong Liu, and Yang Liu · 2023
Later among the works it cites.
Large language models for information retrieval: A survey
Yutao Zhu, Huaying Yuan, Shuting Wang, Jiongnan Liu, Wenhan Liu, Chenlong Deng, Zhicheng Dou, and Ji-Rong Wen · 2023
Later among the works it cites.
Cooperation, competition, and maliciousness: Llm-stakeholders interactive negotiation
Sahar Abdelnabi, Amr Gomaa, Sarath Sivaprasad, Lea Schönherr, and Mario Fritz · 2024
Closest in time.
Playing games with gpt: What can we learn about a large language model from canonical strategic games?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Cited alongside, same era.
Valerio Capraro, Roberto Di Paolo, and Veronica Pizziol · 2023
Cited alongside, same era.
Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models
Neel Guha, Julian Nyarko, Daniel E Ho, Christopher Re, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel Rockmore, et al · 2023
Cited alongside, same era.
Gpt agents in game theory experiments
Fulin Guo · 2023
Cited alongside, same era.
Strategic behavior of large language models: Game structure vs. contextual framing
Babak Heydari and Nunzio Lorè · 2023
Cited alongside, same era.
Large language models as simulated economic agents: What can we learn from homo silicus?
John J Horton · 2023
Cited alongside, same era.
Is chatgpt a good translator? yes with gpt-4 as the engine
Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Xing Wang, Shuming Shi, and Zhaopeng Tu · 2023
Cited alongside, same era.
Assessing the accuracy and reliability of ai-generated medical responses: an evaluation of the chat-gpt model
Douglas Johnson, Rachel Goodman, J Patrinely, Cosby Stone, Eli Zimmerman, Rebecca Donald, Sam Chang, Sean Berkowitz, Avni Finn, Eiman Jahangir, et al · 2023
Cited alongside, same era.
Philip Brookins and Jason DeBacker · 2024
Closest in time.
Put your money where your mouth is: Evaluating strategic planning and execution of llm agents in an auction arena
Jiangjie Chen, Siyu Yuan, Rong Ye, Bodhisattwa Prasad Majumder, and Kyle Richardson · 2024
Closest in time.
Gtbench: Uncovering the strategic reasoning capabilities of llms via game-theoretic evaluations
Jinhao Duan, Renming Zhang, James Diffenderfer, Bhavya Kailkhura, Lichao Sun, Elias Stengel-Eskin, Mohit Bansal, Tianlong Chen, and Kaidi Xu · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
Can large language models serve as rational players in game theory? a systematic analysis
Caoyun Fan, Jindou Chen, Yaohui Jin, and Hao He · 2024
Closest in time.
Suspicion agent: Playing imperfect information games with theory of mind aware gpt-4
Jiaxian Guo, Bo Yang, Paul Yoo, Bill Yuchen Lin, Yusuke Iwasawa, and Yutaka Matsuo · 2024
Closest in time.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al · 2024
Closest in time.
Better zero-shot reasoning with role-play prompting
Aobo Kong, Shiwan Zhao, Hao Chen, Qicheng Li, Yong Qin, Ruiqi Sun, Xin Zhou, Enzhi Wang, and Xiaohang Dong · 2024
Closest in time.
Evaluating large language models in theory of mind tasks
Michal Kosinski · 2024
Closest in time.
Llm-based agent society investigation: Collaboration and confrontation in avalon gameplay
Yihuai Lan, Zhiqiang Hu, Lei Wang, Yang Wang, Deheng Ye, Peilin Zhao, Ee-Peng Lim, Hui Xiong, and Hao Wang · 2024
Closest in time.
Encouraging divergent thinking in large language models through multi-agent debate
Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Zhaopeng Tu, and Shuming Shi · 2024
Closest in time.
Interintent: Investigating social intelligence of llms via intention understanding in an interactive game context
Ziyi Liu, Abhishek Anand, Pei Zhou, Jen-tse Huang, and Jieyu Zhao · 2024
Closest in time.
Our next-generation model: Gemini 1.5
Sundar Pichai and Demis Hassabis · 2024
Closest in time.
Smartplay: A benchmark for llms as intelligent agents
Yue Wu, Xuan Tang, Tom M Mitchell, and Yuanzhi Li · 2024
Closest in time.
Can large language model agents simulate human trust behaviors?
Chengxing Xie, Canyu Chen, Feiran Jia, Ziyu Ye, Shiyang Lai, Kai Shu, Jindong Gu, Adel Bibi, Ziniu Hu, David Jurgens, James Evans, Philip Torr, Bernard Ghanem, and Guohao Li · 2024
Closest in time.
Magic: Investigation of large language model powered multi-agent in cognition, adaptability, rationality and collaboration
Lin Xu, Zhiyuan Hu, Daquan Zhou, Hongyu Ren, Zhen Dong, Kurt Keutzer, See Kiong Ng, and Jiashi Feng · 2024
Closest in time.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al · 2024
Closest in time.
Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu · 2024
Closest in time.
K-level reasoning with large language models
Yadong Zhang, Shaoguang Mao, Tao Ge, Xun Wang, Yan Xia, Man Lan, and Furu Wei · 2024
Closest in time.
Alympics: Llm agents meet game theory
Shaoguang Mao, Yuzhe Cai, Yan Xia, Wenshan Wu, Xun Wang, Fengyi Wang, Qiang Guan, Tao Ge, and Furu Wei · 2025
Closest in time.