Fetching the paper…
Reading the bibliography…
AI Alignment is often presented as an interaction between a single designer and an artificial agent in which the designer attempts to ensure the agent's behavior is consistent with its purpose, and risks arise solely because of conflicts caused by inadvertent misalignment between the utility function intended by the designer and the resulting internal utility function of the agent.
1906
Earlier work this paper cites.
1910
Earlier work this paper cites.
M. Jensen and W. H. Meckling, “Theory of the firm: Managerial behavior, agency costs and ownership structure,” Journal of Financial Economics
1976
Earlier work this paper cites.
D. Cliff, “Evolution of market mechanism through a continuous space of auction-types,” Technical report HPL-2001-326, HP Labs, 2001
2001
Earlier work this paper cites.
J. C. J. M. van den Bergh and S. Stagl, “Coevolution of economic behaviour and institutions: towards a theory of institutional change,” Journal of Evolutionary Economics
2003
Earlier work this paper cites.
2005
Earlier work this paper cites.
K. Alexander, “Corporate governance and banks: The role of regulation in reducing the principal-agent problem,” Journal of Banking regulation
2006
Earlier work this paper cites.
Thesis (phd), University of Liverpool, 2007
S. Phelps, Evolutionary mechanism design · 2007
Earlier work this paper cites.
2009
Earlier work this paper cites.
E. R. Lai, “Metacognition: A literature review,” Always learning: Pearson research report
2011
Earlier work this paper cites.
2017
Cited alongside, same era.
Curran Associates, Inc., 2018
B. Ibarz, J. Leike, T. Pohlen, G. Irving, S. Legg, and D. Amodei, “Reward learning from human preferences and demonstrations in Atari,” in Advances in Neural Information Processing Systems · 2018
Cited alongside, same era.
Oxford University Press, 2021
S. Russell, “Human-compatible artificial intelligence,” in Human-like Machine Intelligence · 2021
Cited alongside, same era.
2022
Cited alongside, same era.
S. Bubeck, V. Chandrasekaran, et al
2023
Cited alongside, same era.
T. Johnson and N. Obradovich, “Evidence of Behavior Consistent with Self-Interest and Altruism in An Artificially Intelligent Agent,” SSRN Electronic Journal
2023
Closest in time.
C. Yang, X. Wang, Y. Lu, H. Liu, Q. V. Le, D. Zhou, and X. Chen, “Large Language Models as Optimizers,” tech. rep., Google DeepMind, 2023 · 2023
Closest in time.
M. Binz and E. Schulz, “Using cognitive psychology to understand gpt-3,” Proceedings of the National Academy of Sciences
2023
Closest in time.
Association for Computing Machinery, 2023
J. S. Park, J. C. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein, Generative Agents: Interactive Simulacra of Human Behavior · 2023
Closest in time.
https://github.com/Significant-Gravitas/Auto-GPT
T. B. Richards, “AutoGTP.” 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Yudkowsky, “Pausing AI Developments Isn’t Enough. We Need to Shut it All Down,” Time Magazine
2023
Cited alongside, same era.
https://github.com/openai/evals
OpenAI, “evals.” 2023 · 2023
Cited alongside, same era.
https://github.com/google/BIG-bench
Google, “BIG-bench.” 2023 · 2023
Cited alongside, same era.
2023
Cited alongside, same era.
Cited in the paper.
OpenAI, “GPT-4 Technical Report,” arXiv:2303.08774
Cited in the paper.
A. Srivastava, A. Rastogi, et al
Cited in the paper.
Closest in time.
M. Tegmark, “The ’Don’t Look Up’ Thinking That Could Doom Us With AI,” Time Magazine
2023
Closest in time.
https://platform.openai.com/docs/guides/chat
OpenAI, “Chat Completion API.” 2023 · 2023
Closest in time.
https://github.com/phelps-sg/llm-cooperation
S. Phelps, 2023 · 2023
Closest in time.
https://github.com/0xk1h0/ChatGPT_DAN
0xk1g0, “ChatGPT DAN.” 2023 · 2023
Closest in time.