Fetching the paper…
Reading the bibliography…
There is increasing interest in using LLMs as decision-making "agents." Doing so includes many degrees of freedom: which model should be used; how should it be prompted; should it be asked to introspect, conduct chain-of-thought reasoning, etc? Settling these questions -- and more broadly, determining whether an LLM agent is reliable enough to be trusted -- requires a methodology for assessing such an agent's economic rationality.
Theory of Games and Economic Behavior
J. von Neumann and O. Morgenstern · 1944
Earlier work this paper cites.
Equilibrium points in n-person games
J. F. Nash Jr · 1950
Earlier work this paper cites.
Cardinal welfare, individualistic ethics, and interpersonal comparisons of utility
J. C. Harsanyi · 1955
Earlier work this paper cites.
Risk, ambiguity, and the savage axioms
D. Ellsberg · 1961
Earlier work this paper cites.
The Foundations of Statistics
L. Savage · 1972
Earlier work this paper cites.
Prospect theory: An analysis of decision under risk
D. Kahneman and A. Tversky · 1979
Earlier work this paper cites.
Sequential equilibria
D. M. Kreps and R. Wilson · 1982
Earlier work this paper cites.
Choices, values, and frames
D. Kahneman and A. Tversky · 1984
Earlier work this paper cites.
Game Theory
D. Fudenberg and J. Tirole · 1991
Earlier work this paper cites.
The strategic implications of sunk costs: A behavioral perspective
R. Parayre · 1995
Earlier work this paper cites.
An introduction to game theory , volume 3
M. J. Osborne et al · 2004
Earlier work this paper cites.
Multiagent systems: Algorithmic, game-theoretic, and logical foundations
Y. Shoham and K. Leyton-Brown · 2008
Earlier work this paper cites.
Decision making under uncertainty: theory and application
M. J. Kochenderfer · 2015
Earlier work this paper cites.
Explanations of the endowment effect: an integrative review
C. K. Morewedge and C. E. Giblin · 2015
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
M. P. Naeini, G. Cooper, and M. Hauskrecht · 2015
Earlier work this paper cites.
From anomalies to forecasts: Toward a descriptive model of decisions under risk, under ambiguity, and from experience
I. Erev, E. Ert, O. Plonsky, D. Cohen, and O. Cohen · 2017
Earlier work this paper cites.
On calibration of modern neural networks
C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger · 2017
Earlier work this paper cites.
Finbert: Financial sentiment analysis with pre-trained language models
D. Araci · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, and J. Brew · 2019
Cited alongside, same era.
HellaSwag: Can a machine really finish your sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Cited alongside, same era.
URL https://openai.com/blog/openai-api
OpenAI, Jun 2020 · 2020
Cited alongside, same era.
Impact of news on the commodity market: Dataset and results
A. Sinha and T. Khandait · 2020
Cited alongside, same era.
FinQA: A dataset of numerical reasoning over financial data
Z. Chen, W. Chen, C. Smiley, S. Shah, I. Borova, D. Langdon, R. Moussa, M. Beane, T.-H. Huang, B. Routledge, and W. Y. Wang · 2021
Cited alongside, same era.
Large language models as simulated economic agents: What can we learn from homo silicus?
J. J. Horton · 2023
Later among the works it cites.
Large language models can self-improve
J. Huang, S. Gu, L. Hou, Y. Wu, X. Wang, H. Yu, and J. Han · 2023
Later among the works it cites.
Agentbench: Evaluating llms as agents, 2023
X. Liu, H. Yu, H. Zhang, Y. Xu, X. Lei, H. Lai, Y. Gu, H. Ding, K. Men, K. Yang, S. Zhang, X. Deng, A. Zeng, Z. Du, C. Zhang, S. Shen, T. Zhang, Y. Su, H. Sun, M. Huang, Y. Dong, and J. Tang · 2023
Later among the works it cites.
Towards accurate differential diagnosis with large language models
D. McDuff, M. Schaekermann, T. Tu, A. Palepu, A. Wang, J. Garrison, K. Singhal, Y. Sharma, S. Azizi, K. Kulkarni, et al · 2023
Later among the works it cites.
Baby agi, 2023
Y. Nakajima · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The GEM benchmark: Natural language generation, its evaluation and metrics
S. Gehrmann, T. Adewumi, K. Aggarwal, P. S. Ammanamanchi, A. Aremu, A. Bosselut, K. R. Chandu, M.-A. Clinciu, D. Das, K. Dhole, W. Du, E. Durmus, O. Dušek, C. C. Emezue, V. Gangal, C. Garbacea, T. Hashimoto, Y. Hou, Y. Jernite, H. Jhamtani, Y. Ji, S. Jolly, M. Kale, D. Kumar, F. Ladhak, A. Madaan, M. Maddela, K. Mahajan, S. Mahamood, B. P. Majumder, P. H. Martins, A. McMillan-Major, S. Mille, E. van Miltenburg, M. Nadeem, S. Narayan, V. Nikolaev, A. Niyongabo Rubungo, S. Osei, A. Parikh, L. Perez-Beltrachini, N. R. Rao, V. Raunak, J. D. Rodriguez, S. Santhanam, J. Sedoc, T. Sellam, S. Shaikh, A. Shimorina, M. A. Sobrevilla Cabezudo, H. Strobelt, N. Subramani, W. Xu, D. Yang, A. Yerukola, and J. Zhou · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2021
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback
Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, et al · 2022
Cited alongside, same era.
Holistic evaluation of language models
P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Kumar, et al · 2022
Cited alongside, same era.
When FLUE meets FLANG: Benchmarks and large pretrained language model for financial domain
R. Shah, K. Chawla, D. Eidnani, A. Shah, W. Du, S. Chava, N. Raman, C. Smiley, J. Chen, and D. Yang · 2022
Cited alongside, same era.
Challenging big-bench tasks and whether chain-of-thought can solve them
M. Suzgun, N. Scales, N. Schärli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. V. Le, E. H. Chi, D. Zhou, , and J. Wei · 2022
Cited alongside, same era.
Chain of thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. H. hsin Chi, F. Xia, Q. Le, and D. Zhou · 2022
Cited alongside, same era.
OpenAI · 2023
Later among the works it cites.
Generative agents: Interactive simulacra of human behavior
J. S. Park, J. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein · 2023
Later among the works it cites.
Tool learning with foundation models, 2023
Y. Qin, S. Hu, Y. Lin, W. Chen, N. Ding, G. Cui, Z. Zeng, Y. Huang, C. Xiao, C. Han, Y. R. Fung, Y. Su, H. Wang, C. Qian, R. Tian, K. Zhu, S. Liang, X. Shen, B. Xu, Z. Zhang, Y. Ye, B. Li, Z. Tang, J. Yi, Y. Zhu, Z. Dai, L. Yan, X. Cong, Y. Lu, W. Zhao, Y. Huang, J. Yan, X. Han, X. Sun, D. Li, J. Phang, C. Yang, T. Wu, H. Ji, Z. Liu, and M. Sun · 2023
Later among the works it cites.
Agentgpt, 2023
Reworkd · 2023
Later among the works it cites.
Toolformer: Language models can teach themselves to use tools
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2023
Later among the works it cites.
Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
Y. Shen, K. Song, X. Tan, D. S. Li, W. Lu, and Y. T. Zhuang · 2023
Later among the works it cites.
Alpaca: A strong, replicable instruction-following model
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Later among the works it cites.
React: Synergizing reasoning and acting in language models, 2023
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao · 2023
Later among the works it cites.
Answering questions by meta-reasoning over multiple chains of thought
O. Yoran, T. Wolfson, B. Bogin, U. Katz, D. Deutch, and J. Berant · 2023
Later among the works it cites.
Webarena: A realistic web environment for building autonomous agents, 2023
S. Zhou, F. F. Xu, H. Zhu, X. Zhou, R. Lo, A. Sridhar, X. Cheng, Y. Bisk, D. Fried, U. Alon, and G. Neubig · 2023
Later among the works it cites.
Ghost in the minecraft: Generally capable agents for open-world environments via large language models with text-based knowledge and memory, 2023
X. Zhu, Y. Chen, H. Tian, C. Tao, W. Su, C. Yang, G. Huang, B. Li, L. Lu, X. Wang, Y. Qiao, Z. Zhang, and J. Dai · 2023
Later among the works it cites.
Mindstorms in natural language-based societies of mind
M. Zhuge, H. Liu, F. Faccio, D. R. Ashley, R. Csordás, A. Gopalakrishnan, A. Hamdi, H. A. A. K. Hammoud, V. Herrmann, K. Irie, et al · 2023
Later among the works it cites.