Fetching the paper…
Reading the bibliography…
An AI control protocol is a plan for usefully deploying AI systems that aims to prevent an AI from intentionally causing some unacceptable outcome.
The strategy of conflict: with a new preface by the author
Schelling, T. C. 1980 · 1980
Earlier work this paper cites.
The absent-minded driver
Aumann, R. J.; Hart, S.; and Perry, M. 1997 · 1997
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P.; Littman, M. L.; and Cassandra, A. R. 1998 · 1998
Earlier work this paper cites.
A survey of point-based POMDP solvers
Shani, G.; Pineau, J.; and Kaplow, R. 2013 · 2013
Earlier work this paper cites.
Measuring Coding Challenge Competence With APPS
Hendrycks, D.; Basart, S.; Kadavath, S.; Mazeika, M.; Arora, A.; Guo, E.; Burns, C.; Puranik, S.; He, H.; Song, D.; and Steinhardt, J. 2021 · 2021
Earlier work this paper cites.
Discovering Language Model Behaviors with Model-Written Evaluations
Perez, E.; Ringer, S.; Lukošiūtė, K.; Nguyen, K.; Chen, E.; Heiner, S.; Pettit, C.; Olsson, C.; Kundu, S.; Kadavath, S.; Jones, A.; Chen, A.; Mann, B.; Israel, B.; Seethor, B.; McKinnon, C.; Olah, C.; Yan, D.; Amodei, D.; Amodei, D.; Drain, D.; Li, D.; Tran-Johnson, E.; Khundadze, G.; Kernion, J.; Landis, J.; Kerr, J.; Mueller, J.; Hyun, J.; Landau, J.; Ndousse, K.; Goldberg, L.; Lovitt, L.; Lucas, M.; Sellitto, M.; Zhang, M.; Kingsland, N.; Elhage, N.; Joseph, N.; Mercado, N.; DasSarma, N.; Rausch, O.; Larson, R.; McCandlish, S.; Johnston, S.; Kravec, S.; Showk, S. E.; Lanham, T.; Telleen-Lawton, T.; Brown, T.; Henighan, T.; Hume, T.; Bai, Y.; Hatfield-Dodds, Z.; Clark, J.; Bowman, S. R.; Askell, A.; Grosse, R.; Hernandez, D.; Ganguli, D.; Hubinger, E.; Schiefer, N.; and Kaplan, J. 2022 · 2022
Earlier work this paper cites.
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
authors, B.-b. 2023 · 2023
Earlier work this paper cites.
Can LLMs Generate Random Numbers? Evaluating LLM Sampling in Controlled Domains
Hopkins, A. K.; and Renda, A. 2023 · 2023
Earlier work this paper cites.
AgentBench: Evaluating LLMs as Agents
Liu, X.; Yu, H.; Zhang, H.; Xu, Y.; Lei, X.; Lai, H.; Gu, Y.; Ding, H.; Men, K.; Yang, K.; Zhang, S.; Deng, X.; Zeng, A.; Du, Z.; Zhang, C.; Shen, S.; Zhang, T.; Su, Y.; Sun, H.; Huang, M.; Dong, Y.; and Tang, J. 2023 · 2023
Earlier work this paper cites.
WildChat: 1M ChatGPT Interaction Logs in the Wild
Zhao, W.; Ren, X.; Hessel, J.; Cardie, C.; Choi, Y.; and Deng, Y. 2023 · 2023
Cited alongside, same era.
Inspect AI: Framework for Large Language Model Evaluations
AI Security Institute, UK. 2024 · 2024
Cited alongside, same era.
Claude Models - Anthropic Documentation
Anthropic. 2024 · 2024
Cited alongside, same era.
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
Costarelli, A.; Allen, M.; Hauksson, R.; Sodunke, G.; Hariharan, S.; Cheng, C.; Li, W.; Clymer, J.; and Yadav, A. 2024 · 2024
Cited alongside, same era.
AI Control: Improving Safety Despite Intentional Subversion
Greenblatt, R.; Shlegeris, B.; Sachan, K.; and Roger, F. 2024 · 2024
Cited alongside, same era.
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs
Mathew, Y.; Matthews, O.; McCarthy, R.; Velja, J.; Witt, C. S. d.; Cope, D.; and Schoots, N. 2024 · 2024
Closest in time.
Secret Collusion Among Generative AI Agents
Motwani, S. R.; Baranchuk, M.; Strohmeier, M.; Bolina, V.; Torr, P. H. S.; Hammond, L.; and de Witt, C. S. 2024 · 2024
Closest in time.
OpenAI o1 System Card
OpenAI. 2024 · 2024
Closest in time.
Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
Wen, J.; Hebbar, V.; Larson, C.; Bhatt, A.; Radhakrishnan, A.; Sharma, M.; Sleight, H.; Feng, S.; He, H.; Perez, E.; Shlegeris, B.; and Khan, A. 2024 · 2024
Closest in time.
Ctrl-Z: Controlling AI Agents via Resampling (forthcoming)
Bhatt, A.; Rushing, C.; Kaufman, A.; Tracy, T.; Georgiev, V.; Khan, A.; Matolcsi, D.; and Shlegeris, B. 2025 · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Griffin, C.; Thomson, L.; Shlegeris, B.; and Abate, A. 2024 · 2024
Cited alongside, same era.
Me, myself, and AI: The situational awareness dataset (SAD) for LLMs
Laine, R.; Chughtai, B.; Betley, J.; Hariharan, K.; Balesni, M.; Scheurer, J.; Hobbhahn, M.; Meinke, A.; and Evans, O. 2024 · 2024
Cited alongside, same era.
The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
Li, N.; Pan, A.; Gopal, A.; Yue, S.; Berrios, D.; Gatti, A.; Li, J. D.; Dombrowski, A.-K.; Goel, S.; Phan, L.; Mukobi, G.; Helm-Burger, N.; Lababidi, R.; Justen, L.; Liu, A. B.; Chen, M.; Barrass, I.; Zhang, O.; Zhu, X.; Tamirisa, R.; Bharathi, B.; Khoja, A.; Zhao, Z.; Herbert-Voss, A.; Breuer, C. B.; Marks, S.; Patel, O.; Zou, A.; Mazeika, M.; Wang, Z.; Oswal, P.; Lin, W.; Hunt, A. A.; Tienken-Harder, J.; Shih, K. Y.; Talley, K.; Guan, J.; Kaplan, R.; Steneker, I.; Campbell, D.; Jokubaitis, B.; Levinson, A.; Wang, J.; Qian, W.; Karmakar, K. K.; Basart, S.; Fitz, S.; Levine, M.; Kumaraguru, P.; Tupakula, U.; Varadharajan, V.; Wang, R.; Shoshitaishvili, Y.; Ba, J.; Esvelt, K. M.; Wang, A.; and Hendrycks, D. 2024 · 2024
Cited alongside, same era.
Subversion Strategy Eval: Evaluating AI’s stateless strategic capabilities against control protocols
Mallen, A.; Griffin, C.; Abate, A.; and Shlegeris, B. 2024 · 2024
Cited alongside, same era.
Closest in time.
OpenAI o3-mini System Card
OpenAI. 2025 · 2025
Closest in time.
Thoughts on the conservative assumptions in AI control
Shlegeris, B. 2025 · 2025
Closest in time.
Estimating the Probabilities of Rare Outputs in Language Models
Wu, G.; and Hilton, J. 2025 · 2025
Closest in time.