Fetching the paper…
Reading the bibliography…
We develop evaluation methods for measuring the economic decision-making capabilities and tendencies of LLMs.
College Admissions and the Stability of Marriage
D. Gale and L. S. Shapley · 1962
Earlier work this paper cites.
Marriages stables
Donald E. Knuth · 1976
Earlier work this paper cites.
Estimating Discrete-Choice Models of Product Differentiation
Steven T. Berry · 1994
Earlier work this paper cites.
Are emily and greg more employable than lakisha and jamal? a field experiment on labor market discrimination
Marianne Bertrand and Sendhil Mullainathan · 2004
Earlier work this paper cites.
On the complexity of trial and error
Xiaohui Bei, Ning Chen, and Shengyu Zhang · 2013
Earlier work this paper cites.
Global evidence on economic preferences*
Armin Falk, Anke Becker, Thomas Dohmen, Benjamin Enke, David Huffman, and Uwe Sunde · 2018
Earlier work this paper cites.
Developing competition law for collusion by autonomous artificial agents
Joseph E Harrington, Jr · 2018
Earlier work this paper cites.
The Complexity of Interactively Learning a Stable Matching by Trial and Error
Ehsan Emamjomeh-Zadeh, Yannai A. Gonczarowski, and David Kempe · 2020
Earlier work this paper cites.
On the fairness of machine-assisted human decisions, 2021
Talia Gillis, Bryce McLaughlin, and Jann Spiess · 2021
Earlier work this paper cites.
Testing the waters: Behavior across participant pools
Erik Snowberg and Leeat Yariv · 2021
Earlier work this paper cites.
Artificial intelligence and auction design
Martino Banchio and Andrzej Skrzypacz · 2022
Earlier work this paper cites.
Algorithmic design: Fairness versus accuracy
Annie Liang, Jay Lu, and Xiaosheng Mu · 2022
Earlier work this paper cites.
Using large language models to simulate multiple humans and replicate human subject studies
Gati Aher, Rosa I. Arriaga, and Adam Tauman Kalai · 2023
Earlier work this paper cites.
Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?, April 2023
John J. Horton · 2023
Earlier work this paper cites.
Frontiers: Can Large Language Models Capture Human Preferences?
Ali Goli and Amandeep Singh · 2023
Earlier work this paper cites.
Adaptive algorithms and collusion via coupling
Martino Banchio and Giacomo Mantegazza · 2023
Earlier work this paper cites.
Inverse selection, 2023
Markus K Brunnermeier, Rohit Lamba, and Carlos Segura-Rodriguez · 2023
Earlier work this paper cites.
Generative AI at Work, April 2023
Erik Brynjolfsson, Danielle Li, and Lindsey R. Raymond · 2023
Earlier work this paper cites.
The market effects of algorithms
Lindsey Raymond · 2023
Earlier work this paper cites.
Adversarial competition and collusion in algorithmic markets
Luc Rocher, Arnaud J Tournier, and Yves-Alexandre de Montjoye · 2023
Earlier work this paper cites.
FinanceBench: A New Benchmark for Financial Question Answering, November 2023
Pranab Islam, Anand Kannappan, Douwe Kiela, Rebecca Qian, Nino Scherrer, and Bertie Vidgen · 2023
Earlier work this paper cites.
Voyager: An Open-Ended Embodied Agent with Large Language Models, October 2023
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar · 2023
Earlier work this paper cites.
GAIA: a benchmark for General AI Assistants
Grégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun, and Thomas Scialom · 2023
Earlier work this paper cites.
WebArena: A Realistic Web Environment for Building Autonomous Agents
Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig · 2023
Earlier work this paper cites.
AgentBench: Evaluating LLMs as Agents
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang · 2023
Earlier work this paper cites.
Karthik Valmeekam, Matthew Marquez, Alberto Olmo, Sarath Sreedharan, and Subbarao Kambhampati · 2023
Cited alongside, same era.
Alexander Pan, Jun Shern Chan, Andy Zou, Nathaniel Li, Steven Basart, Thomas Woodside, Hanlin Zhang, Scott Emmons, and Dan Hendrycks · 2023
Cited alongside, same era.
Toolformer: Language Models Can Teach Themselves to Use Tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessi, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom · 2023
Cited alongside, same era.
Welfare Distribution in Two-sided Random Matching Markets
Itai Ashlagi, Mark Braverman, and Geng Zhao · 2023
Cited alongside, same era.
Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr · 2024
Later among the works it cites.
ProSA: Assessing and understanding the prompt sensitivity of LLMs
Jingming Zhuo, Songyang Zhang, Xinyu Fang, Haodong Duan, Dahua Lin, and Kai Chen · 2024
Later among the works it cites.
Regulation of algorithmic collusion
Jason D. Hartline, Sheng Long, and Chenhao Zhang · 2024
Later among the works it cites.
Algorithmic collusion: where are we and where should we be going?, August 2024
Ibrahim Abada, Joseph E Harrington, Jr, Xavier Lambin, and Janusz M Meylahn · 2024
Later among the works it cites.
AI capabilities have steadily improved over the past year
Luke Emberson · 2025
Closest in time.
Visa Launches AI Agents for Shopping, May 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Benjamin S. Manning, Kehang Zhu, and John J. Horton · 2024
Cited alongside, same era.
Can large language models explore in-context?
Akshay Krishnamurthy, Keegan Harris, Dylan J. Foster, Cyril Zhang, and Aleksandrs Slivkins · 2024
Cited alongside, same era.
LLMs at the Bargaining Table
Yuan Deng, Vahab Mirrokni, Renato Paes Leme, Hanrui Zhang, and Song Zuo · 2024
Cited alongside, same era.
Algorithmic Collusion by Large Language Models, November 2024
Sara Fish, Yannai A. Gonczarowski, and Ran I. Shorrer · 2024
Cited alongside, same era.
EconLogicQA: A Question-Answering Benchmark for Evaluating Large Language Models in Economic Sequential Reasoning
Yinzhu Quan and Zefang Liu · 2024
Cited alongside, same era.
STEER: assessing the economic rationality of large language models
Narun Raman, Taylor Lundy, Samuel Joseph Amouyal, Yoav Levine, Kevin Leyton-Brown, and Moshe Tennenholtz · 2024
Cited alongside, same era.
TheAgentCompany: Benchmarking LLM agents on consequential real world tasks, 2024
Frank F. Xu, Yufan Song, Boxuan Li, Yuxuan Tang, Kritanjali Jain, Mengxue Bao, Zora Z. Wang, Xuhui Zhou, Zhitong Guo, Murong Cao, Mingyang Yang, Hao Yang Lu, Amaad Martin, Zhe Su, Leander Maben, Raj Mehta, Wayne Chi, Lawrence Jang, Yiqing Xie, Shuyan Zhou, and Graham Neubig · 2024
Cited alongside, same era.
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
Chang Ma, Junlei Zhang, Zhihao Zhu, Cheng Yang, Yujiu Yang, Yaohui Jin, Zhenzhong Lan, Lingpeng Kong, and Junxian He · 2024
Cited alongside, same era.
Bloomberg · 2025
Closest in time.
How Generative AI Improves Supply Chain Management
Ishai Menache, Jeevan Pathuri, David Simchi-Levi, and Tom Linton · 2025
Closest in time.
Which economic tasks are performed with ai? evidence from millions of claude conversations, 2025
Kunal Handa, Alex Tamkin, Miles McCain, Saffron Huang, Esin Durmus, Sarah Heck, Jared Mueller, Jerry Hong, Stuart Ritchie, Tim Belonax, Kevin K. Troy, Dario Amodei, Jared Kaplan, Jack Clark, and Deep Ganguli · 2025
Closest in time.
Measuring Perceived Slant in Large Language Models Through User Evaluations, May 2025
Sean J Westwood, Justin Grimmer, and Andrew B Hall · 2025
Closest in time.
Steer-me: Assessing the microeconomic reasoning of large language models
Narun Raman, Taylor Lundy, Thiago Amin, Kevin Leyton-Brown, and Jesse Perla · 2025
Closest in time.
Econwebarena: Benchmarking autonomous agents on economic tasks in realistic web environments
Zefang Liu and Yinzhu Quan · 2025
Closest in time.
Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents, February 2025
Axel Backlund and Lukas Petersson · 2025
Closest in time.
Humanity’s Last Exam, January 2025
Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Sean Shi, Michael Choi, Anish Agrawal, Arnav Chopra, Adam Khoja, Richard Ren, Jason Hausenloy, Oliver Zhang, Mantas Mazeika, Summer Yue, Alexandr Wang, and Dan Hendrycks · 2025
Closest in time.
GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks, October 2025
Tejal Patwardhan, Rachel Dias, Elizabeth Proehl, Grace Kim, Michele Wang, Olivia Watkins, Simón Posada Fishman, Marwan Aljubeh, Phoebe Thacker, Laurance Fauconnet, Natalie S. Kim, Patrick Chao, Samuel Miserendino, Gildas Chabot, David Li, Michael Sharman, Alexandra Barr, Amelia Glaese, and Jerry Tworek · 2025
Closest in time.
ARC Prize 2024: Technical Report, January 2025
Francois Chollet, Mike Knoop, Gregory Kamradt, and Bryan Landers · 2025
Closest in time.
Utility engineering: Analyzing and controlling emergent value systems in AIs, 2025
Mantas Mazeika, Xuwang Yin, Rishub Tamirisa, Jaehyuk Lim, Bruce W. Lee, Richard Ren, Long Phan, Norman Mu, Adam Khoja, Oliver Zhang, and Dan Hendrycks · 2025
Closest in time.
Playing repeated games with large language models
Elif Akata, Lion Schulz, Julian Coda-Forno, Seong Joon Oh, Matthias Bethge, and Eric Schulz · 2025
Closest in time.
Jen-tse Huang, Eric John Li, Man Ho Lam, Tian Liang, Wenxuan Wang, Youliang Yuan, Wenxiang Jiao, Xing Wang, Zhaopeng Tu, and Michael R. Lyu · 2025
Closest in time.
LLM-Coordination: Evaluating and Analyzing Multi-agent Coordination Abilities in Large Language Models
Saaket Agashe, Yue Fan, Anthony Reyna, and Xin Eric Wang · 2025
Closest in time.
Using cognitive models to reveal value trade-offs in language models, October 2025
Sonia K. Murthy, Rosie Zhao, Jennifer Hu, Sham Kakade, Markus Wulfmeier, Peng Qian, and Tomer Ullman · 2025
Closest in time.
Infrastructure for AI agents, 2025
Alan Chan, Kevin Wei, Sihao Huang, Nitarshan Rajkumar, Elija Perrier, Seth Lazar, Gillian K. Hadfield, and Markus Anderljung · 2025
Closest in time.
Measuring AI ability to complete long tasks
METR · 2025
Closest in time.
Human Misperception of Generative-AI Alignment: A Laboratory Experiment
Kevin He, Ran Shorrer, and Mengjia Xia · 2025
Closest in time.
Artificial Intelligence, Algorithmic Pricing, and Collusion
Emilio Calvano, Giacomo Calzolari, Vincenzo Denicolò, and Sergio Pastorello · 2025
Closest in time.
JPMorgan Replaces Proxy Advisers With AI for Voting US Shares
Jennifer Surane · 2026
Closest in time.