Fetching the paper…
Reading the bibliography…
As AI systems pervade human life, ensuring that large language models (LLMs) make safe decisions remains a significant challenge.
Variabilità e mutabilità: contributo allo studio delle distribuzioni e delle relazioni statistiche.[Fasc. I.]
C. Gini · 1912
Earlier work this paper cites.
The economic theory of a common-property resource: the fishery
H. S. Gordon · 1954
Earlier work this paper cites.
The tragedy of the commons
G. Hardin · 1968
Earlier work this paper cites.
The evolution of cooperation
R. Axelrod and W. D. Hamilton · 1981
Earlier work this paper cites.
Governing the commons: The evolution of institutions for collective action
E. Ostrom · 1990
Earlier work this paper cites.
Order without law: How neighbors settle disputes
R. C. Ellickson · 1991
Earlier work this paper cites.
Revisiting the commons: local lessons, global challenges
E. Ostrom, J. Burger, C. B. Field, R. B. Norgaard, and D. Policansky · 1999
Earlier work this paper cites.
Aligning ai with shared human values
D. Hendrycks, C. Burns, S. Basart, A. Critch, J. Li, D. Song, and J. Steinhardt · 2008
Earlier work this paper cites.
Multiagent systems: Algorithmic, game-theoretic, and logical foundations
Y. Shoham and K. Leyton-Brown · 2008
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2009
Earlier work this paper cites.
Behavioral game theory: Experiments in strategic interaction
C. F. Camerer · 2011
Earlier work this paper cites.
Human cooperation
D. G. Rand and M. A. Nowak · 2013
Earlier work this paper cites.
Origins of human cooperation and morality
M. Tomasello and A. Vaish · 2013
Earlier work this paper cites.
Moral tribes: Emotion, reason, and the gap between us and them
J. Greene · 2014
Earlier work this paper cites.
Researchers warn against ’autonomous weapons’ arms race, 2020
NPR · 2015
Earlier work this paper cites.
Concrete problems in ai safety
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Earlier work this paper cites.
Coordinate to cooperate or compete: abstract goals and joint intentions in social interaction
M. Kleiman-Weiner, M. K. Ho, J. L. Austerweil, M. L. Littman, and J. B. Tenenbaum · 2016
Earlier work this paper cites.
A multi-agent reinforcement learning model of common-pool resource appropriation
J. Perolat, J. Z. Leibo, V. Zambaldi, C. Beattie, K. Tuyls, and T. Graepel · 2017
Earlier work this paper cites.
Life 3.0: Being Human in the Age of Artificial Intelligence
M. Tegmark · 2017
Earlier work this paper cites.
Finding friend and foe in multi-agent games
J. Serrino, M. Kleiman-Weiner, D. C. Parkes, and J. Tenenbaum · 2019
Earlier work this paper cites.
Theory of minds: Understanding behavior in groups through inverse planning
M. Shum, M. Kleiman-Weiner, M. L. Littman, and J. B. Tenenbaum · 2019
Earlier work this paper cites.
Ai research considerations for human existential safety (arches)
A. Critch and D. Krueger · 2020
Earlier work this paper cites.
Open problems in cooperative ai
A. Dafoe, E. Hughes, Y. Bachrach, T. Collins, K. R. McKee, J. Z. Leibo, K. Larson, and T. Graepel · 2020
Earlier work this paper cites.
The logic of universalization guides moral judgment
S. Levine, M. Kleiman-Weiner, L. Schulz, J. Tenenbaum, and F. Cushman · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Cited alongside, same era.
Cooperative ai: machines must learn to find common ground
A. Dafoe, Y. Bachrach, G. Hadfield, E. Horvitz, K. Larson, and T. Graepel · 2021
Cited alongside, same era.
Are we learning yet? a meta review of evaluation failures across machine learning
T. Liao, R. Taori, I. D. Raji, and L. Schmidt · 2021
Cited alongside, same era.
When to make exceptions: Exploring language models as accounts of human moral judgment
Z. Jin, S. Levine, F. Gonzalez Adauto, O. Kamal, M. Sap, M. Sachan, R. Mihalcea, J. Tenenbaum, and B. Schölkopf · 2022
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods, 2022
S. Lin, J. Hilton, and O. Evans · 2022
Cited alongside, same era.
Social simulacra: Creating populated prototypes for social computing systems
Agentsims: An open-source sandbox for large language model evaluation
J. Lin, H. Zhao, A. Zhang, Y. Wu, H. Ping, and Q. Chen · 2023
Later among the works it cites.
How do we know how smart ai systems are?, 2023
M. Mitchell · 2023
Later among the works it cites.
Dera: enhancing large language model completions with dialog-enabled resolving agents
V. Nair, E. Schumacher, G. Tso, and A. Kannan · 2023
Later among the works it cites.
World models for math story problems
A. Opedal, N. Stoehr, A. Saparov, and M. Sachan · 2023
Later among the works it cites.
Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark
A. Pan, J. S. Chan, A. Zou, N. Li, S. Basart, T. Woodside, J. Ng, H. Zhang, S. Emmons, and D. Hendrycks · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. S. Park, L. Popowski, C. Cai, M. R. Morris, P. Liang, and M. S. Bernstein · 2022
Cited alongside, same era.
The ai economist: Taxation policy design via two-level deep multiagent reinforcement learning
S. Zheng, A. Trott, S. Srinivasa, D. C. Parkes, and R. Socher · 2022
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, et al · 2023
Cited alongside, same era.
Managing ai risks in an era of rapid progress
Y. Bengio, G. Hinton, A. Yao, D. Song, P. Abbeel, Y. N. Harari, Y.-Q. Zhang, L. Xue, S. Shalev-Shwartz, G. Hadfield, et al · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with GPT-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. M. Lundberg, H. Nori, H. Palangi, M. T. Ribeiro, and Y. Zhang · 2023
Cited alongside, same era.
Foundations of cooperative ai
V. Conitzer and C. Oesterheld · 2023
Cited alongside, same era.
Later among the works it cites.
Generative agents: Interactive simulacra of human behavior
J. S. Park, J. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
Later among the works it cites.
Humanoid agents: Platform for simulating human-like generative agents
Z. Wang, Y. Y. Chiu, and Y. C. Chiu · 2023
Later among the works it cites.
The rise and potential of large language model based agents: A survey, 2023
Z. Xi, W. Chen, X. Guo, W. He, Y. Ding, B. Hong, M. Zhang, J. Wang, S. Jin, E. Zhou, R. Zheng, X. Fan, X. Wang, L. Xiong, Y. Zhou, W. Wang, C. Jiang, Y. Zou, X. Liu, Z. Yin, S. Dou, R. Weng, W. Cheng, Q. Zhang, W. Qin, Y. Zheng, X. Qiu, X. Huang, and T. Gui · 2023
Later among the works it cites.
Exploring large language models for communication games: An empirical study on werewolf
Y. Xu, S. Wang, P. Li, F. Luo, X. Wang, W. Liu, and Y. Liu · 2023
Later among the works it cites.
Exploring collaboration mechanisms for llm agents: A social psychology view
J. Zhang, X. Xu, and S. Deng · 2023
Later among the works it cites.
Webarena: A realistic web environment for building autonomous agents
S. Zhou, F. F. Xu, H. Zhu, X. Zhou, R. Lo, A. Sridhar, X. Cheng, Y. Bisk, D. Fried, U. Alon, et al · 2023
Later among the works it cites.
Foundational challenges in assuring alignment and safety of large language models
U. Anwar, A. Saparov, J. Rando, D. Paleka, M. Turpin, P. Hase, E. S. Lubana, E. Jenner, S. Casper, O. Sourbut, et al · 2024
Closest in time.
When is it acceptable to break the rules? knowledge representation of moral judgements based on empirical data
E. Awad, S. Levine, A. Loreggia, N. Mattei, I. Rahwan, F. Rossi, K. Talamadupula, J. Tenenbaum, and M. Kleiman-Weiner · 2024
Closest in time.
Chatbot arena: An open platform for evaluating llms by human preference
W.-L. Chiang, L. Zheng, Y. Sheng, A. N. Angelopoulos, T. Li, D. Li, H. Zhang, B. Zhu, M. Jordan, J. E. Gonzalez, et al · 2024
Closest in time.
URL https://www.cognition-labs.com/introducing-devin
Cognition, 2024 · 2024
Closest in time.
Evaluating language model agency through negotiations
T. R. Davidson, V. Veselovsky, M. Josifoski, M. Peyrard, A. Bosselut, M. Kosinski, and R. West · 2024
Closest in time.
Mind2web: Towards a generalist agent for the web
X. Deng, Y. Gu, B. Zheng, S. Chen, S. Stevens, B. Wang, H. Sun, and Y. Su · 2024
Closest in time.
Gtbench: Uncovering the strategic reasoning limitations of llms via game-theoretic evaluations
J. Duan, R. Zhang, J. Diffenderfer, B. Kailkhura, L. Sun, E. Stengel-Eskin, M. Bansal, T. Chen, and K. Xu · 2024
Closest in time.
When rules are over-ruled: Virtual bargaining as a contractualist method of moral judgment
S. Levine, M. Kleiman-Weiner, N. Chater, F. Cushman, and J. B. Tenenbaum · 2024
Closest in time.
Camel: Communicative agents for" mind" exploration of large language model society
G. Li, H. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem · 2024
Closest in time.
Steer: Assessing the economic rationality of large language models
N. Raman, T. Lundy, S. Amouyal, Y. Levine, K. Leyton-Brown, and M. Tennenholtz · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al · 2024
Closest in time.