Fetching the paper…
Reading the bibliography…
As LLM agents gain a greater capacity to cause harm, AI developers might increasingly rely on control measures such as monitoring to justify that they are safe.
Failure mode and effects analysis (FMEA): A guide for continuous improvement for the semiconductor equipment industry
Mario Villacourt · 1992
Earlier work this paper cites.
Defence standard 00-56 issue 4: Safety management requirements for defence systems
UK Ministry of Defence · 2007
Earlier work this paper cites.
Has the safety case failed?
Brendan Fitzgerald, Paul Breen, and Joe Patrick · 2010
Earlier work this paper cites.
Software certification: Is there a case against safety cases?
Alan Wassyng, Tom Maibaum, Mark Lawford, and Hans Bherer · 2011
Earlier work this paper cites.
Building blocks for assurance cases
Robin Bloomfield and Kateryna Netkachova · 2014
Earlier work this paper cites.
ISO 9001:2015 quality management systems – requirements
International Organization for Standardization · 2015
Earlier work this paper cites.
Geoffrey Irving, Paul Christiano, and Dario Amodei · 2018
Earlier work this paper cites.
STPA Handbook
Nancy G. Leveson and John P. Thomas · 2018
Earlier work this paper cites.
Modelling confidence in railway safety case
Rui Wang, Jérémie Guiochet, Gilles Motet, and Walter Schön · 2018
Earlier work this paper cites.
Safety cases: Past, present and future
Ewen Denney and Ganesh Pai · 2019
Earlier work this paper cites.
Implementation of nuclear safety cases
Alec Bounds · 2020
Earlier work this paper cites.
Safety cases in the certification of autonomous systems
Thor Myklebust and Tor Stålhane · 2020
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan · 2022
Earlier work this paper cites.
Taken out of context: On measuring situational awareness in llms, 2023
Lukas Berglund, Asa Cooper Stickland, Mikita Balesni, Max Kaufmann, Meg Tong, Tomasz Korbak, Daniel Kokotajlo, and Owain Evans · 2023
Earlier work this paper cites.
Scheming AIs: Will AIs fake alignment during training in order to get power?, 2023
Joe Carlsmith · 2023
Earlier work this paper cites.
Auditing failures vs concentrated failures, December 2023
Ryan Greenblatt and Fabien Roger · 2023
Earlier work this paper cites.
AI control: Improving safety despite intentional subversion
Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger · 2023
Earlier work this paper cites.
Pretraining language models with human preferences
Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Bhalerao, Christopher L. Buckley, Jason Phang, Samuel R. Bowman, and Ethan Perez · 2023
Earlier work this paper cites.
Preventing language models from hiding their reasoning, 2023
Fabien Roger and Ryan Greenblatt · 2023
Cited alongside, same era.
Reflexion: Language agents with verbal reinforcement learning, 2023
Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2023
Cited alongside, same era.
React: Synergizing reasoning and acting in language models, 2023
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao · 2023
Cited alongside, same era.
International scientific report on the safety of advanced AI, May 2024
AI Seoul Summit Scientific Advisory Group · 2024
Cited alongside, same era.
Responsible scaling policy
Anthropic · 2024
Cited alongside, same era.
Three sketches of ASL-4 safety case components
Roger Grosse · 2024
Later among the works it cites.
Safety cases at AISI
Geoffrey Irving · 2024
Later among the works it cites.
SWE-bench: Can language models resolve real-world GitHub issues?
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan · 2024
Later among the works it cites.
Risk thresholds for frontier AI, 2024
Leonie Koessler, Jonas Schuett, and Markus Anderljung · 2024
Later among the works it cites.
Me, myself, and AI: The situational awareness dataset (SAD) for LLMs
Rudolf Laine, Bilal Chughtai, Jan Betley, Kaivalya Hariharan, Mikita Balesni, Jérémy Scheurer, Marius Hobbhahn, Alexander Meinke, and Owain Evans · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards evaluations-based safety cases for AI scheming, 2024
Mikita Balesni, Marius Hobbhahn, David Lindner, Alexander Meinke, Tomek Korbak, Joshua Clymer, Buck Shlegeris, Jérémy Scheurer, Charlotte Stix, Rusheb Shah, Nicholas Goldowsky-Dill, Dan Braun, Bilal Chughtai, Owain Evans, Daniel Kokotajlo, and Lucius Bushnaq · 2024
Cited alongside, same era.
Managing extreme AI risks amid rapid progress
Yoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Trevor Darrell, Yuval Noah Harari, Ya-Qin Zhang, Lan Xue, Shai Shalev-Shwartz, Gillian Hadfield, Jeff Clune, Tegan Maharaj, Frank Hutter, Atılım Güneş Baydin, Sheila McIlraith, Qiqi Gao, Ashwin Acharya, David Krueger, Anca Dragan, Philip Torr, Stuart Russell, Daniel Kahneman, Jan Brauner, and Sören Mindermann · 2024
Cited alongside, same era.
Sabotage evaluations for frontier models
Joe Benton, Misha Wagner, Eric Christiansen, Cem Anil, Ethan Perez, Jai Srivastav, Esin Durmus, Deep Ganguli, Shauna Kravec, Buck Shlegeris, Jared Kaplan, Holden Karnofsky, Evan Hubinger, Roger Grosse, Samuel R. Bowman, and David Duvenaud · 2024
Cited alongside, same era.
Shell games: Control protocols for adversarial AI agents, 2024
Aryan Bhatt, Cody Rushing, Adam Kaufman, Vasil Georgiev, Tyler Tracy, Akbir Khan, and Buck Shlegeris · 2024
Cited alongside, same era.
Looking inward: Language models can learn about themselves by introspection, 2024
Felix J Binder, James Chua, Tomek Korbak, Henry Sleight, John Hughes, Robert Long, Ethan Perez, Miles Turpin, and Owain Evans · 2024
Cited alongside, same era.
The checklist: What succeeding at AI safety will involve
Sam Bowman · 2024
Cited alongside, same era.
Safety cases for frontier AI, 2024
Marie Davidsen Buhl, Gaurav Sett, Leonie Koessler, Jonas Schuett, and Markus Anderljung · 2024
Cited alongside, same era.
Jan Leike, John Schulman, and Jeffrey Wu · 2024
Later among the works it cites.
LLM critics help catch LLM bugs, 2024
Nat McAleese, Rai Michael Pokorny, Juan Felipe Ceron Uribe, Evgenia Nitishinskaya, Maja Trebacz, and Jan Leike · 2024
Later among the works it cites.
Frontier models are capable of in-context scheming, 2024
Alexander Meinke, Bronson Schoen, Jérémy Scheurer, Mikita Balesni, Rusheb Shah, and Marius Hobbhahn · 2024
Later among the works it cites.
The operational risks of AI in large-scale biological attacks: Results of a red-team study, 2024
Christopher A. Mouton, Caleb Lucas, and Ella Guest · 2024
Later among the works it cites.
NIST AI RMF 4.3: Incidents and errors are communicated to relevant AI actors
National Institute of Standards and Technology · 2024
Later among the works it cites.
Evaluating frontier models for dangerous capabilities
Mary Phuong, Matthew Aitchison, Elliot Catt, Sarah Cogan, Alexandre Kaskasoli, Victoria Krakovna, David Lindner, Matthew Rahtz, et al · 2024
Later among the works it cites.
GPQA: A graduate-level google-proof q&a benchmark
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman · 2024
Later among the works it cites.
Ai sandbagging: Language models can strategically underperform on evaluations, 2024
Teun van der Weij, Felix Hofstätter, Ollie Jaffe, Samuel F. Brown, and Francis Rhys Ward · 2024
Later among the works it cites.
Adaptive deployment of untrusted llms reduces distributed threats, 2024
Jiaxin Wen, Vivek Hebbar, Caleb Larson, Aryan Bhatt, Ansh Radhakrishnan, Mrinank Sharma, Henry Sleight, Shi Feng, He He, Ethan Perez, Buck Shlegeris, and Akbir Khan · 2024
Later among the works it cites.
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Hjalmar Wijk, Tao Lin, Joel Becker, Sami Jawhar, Neev Parikh, Thomas Broadley, Lawrence Chan, Michael Chen, Josh Clymer, Jai Dhyani, Elena Ericheva, Katharyn Garcia, Brian Goodrich, Nikola Jurkovic, Megan Kinniment, Aron Lajko, Seraphina Nix, Lucas Sato, William Saunders, Maksym Taran, Ben West, and Elizabeth Barnes · 2024
Later among the works it cites.
Fine-tuning language models from human preferences, 2020
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2024
Later among the works it cites.
Extending control evaluations to non-scheming threats
Josh Clymer · 2025
Closest in time.
Thoughts on the conservative assumptions in AI control
Buck Shlegeris · 2025
Closest in time.