Fetching the paper…
Reading the bibliography…
Successful self-replication under no human assistance is the essential step for AI to outsmart the human beings, and is an early signal for rogue AIs.
Theory of Self Reproducing Automata (University of Illinois Press, 1966)
von Neumann, J. & Burks, A. W · 1966
Earlier work this paper cites.
Life 3.0: Being human in the age of artificial intelligence (Vintage, 2018)
Tegmark, M · 2018
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J. et al · 2020
Earlier work this paper cites.
Discovering Language Model Behaviors with Model-Written Evaluations
Perez, E. et al · 2022
Earlier work this paper cites.
Anthropic’s Responsible Scaling Policy (2023)
Anthropic · 2023
Earlier work this paper cites.
Evaluating language-model agents on realistic autonomous tasks
Kinniment, M. et al · 2023
Earlier work this paper cites.
Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection
Greshake, K. et al · 2023
Earlier work this paper cites.
Autonomous replication and adaptation: an attempt at a concrete danger threshold (2023)
Hjalmar Wijk · 2023
Cited alongside, same era.
Asilomar ai principles (2017)
The Beneficial AI 2017 Conference · 2024
Cited alongside, same era.
OpenAI’s Safety Policy (2024)
OpenAI · 2024
Cited alongside, same era.
Google DeepMind’s Frontier Safety Framework (2024)
Google DeepMind · 2024
Cited alongside, same era.
Openai’s preparedness framework (2023)
OpenAI · 2024
Cited alongside, same era.
Evaluating frontier models for dangerous capabilities
Phuong, M. et al · 2024
Cited alongside, same era.
OpenAI o1 System Card (New)
OpenAI · 2024
Closest in time.
Meta’s llama 3.1 (2024)
Meta Inc · 2024
Closest in time.
Qwen2.5: A Party of Foundation Models! (2024)
Alibaba Inc · 2024
Closest in time.
Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions
Zhuo, T. Y. et al · 2024
Closest in time.
Chatbot arena: An open platform for evaluating llms by human preference
Chiang, W.-L. et al · 2024
Closest in time.
The shutdown problem: an AI engineering puzzle for decision theorists
Thornley, E · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
OpenAI · 2024
Cited alongside, same era.