Fetching the paper…
Reading the bibliography…
Over the past year, artificial intelligence (AI) companies have been increasingly adopting AI safety frameworks.
The contribution of latent human failures to the breakdown of complex systems
J. Reason · 1990
Earlier work this paper cites.
Human error
J. Reason · 1990
Earlier work this paper cites.
Deterministic versus probabilistic risk assessment: Strengths and weaknesses in a regulatory context
G. M. Richardson · 1996
Earlier work this paper cites.
On the use of probabilistic and deterministic methods in risk analysis
C. Kirchsteiger · 1999
Earlier work this paper cites.
Normal accidents: Living with high-risk technologies
C. Perrow · 2000
Earlier work this paper cites.
A risk informed defense-in-depth framework for existing and advanced reactors
K. N. Fleming and F. A. Silady · 2002
Earlier work this paper cites.
Expert judgement elicitation for risk assessments of critical infrastructures
R. M. Cooke and L. H. J. Goossens · 2004
Earlier work this paper cites.
Considering defense in depth for software applications
M. R. Stytz · 2004
Earlier work this paper cites.
The Delphi technique: Making sense of consensus
C.-C. Hsu and B. A. Sandford · 2007
Earlier work this paper cites.
The black swan: The impact of the highly improbable
N. N. Taleb · 2007
Earlier work this paper cites.
A route to more tractable expert advice
W. Aspinall · 2010
Earlier work this paper cites.
Integrated aviation security for defense-in-depth of next generation air transportation system
W. Li and P. Kamal · 2011
Earlier work this paper cites.
On the meaning of a black swan in a risk context
T. Aven · 2013
Earlier work this paper cites.
IIA position paper: The Three Lines of Defense in effective risk management and control
Institute of Internal Auditors · 2013
Earlier work this paper cites.
Implementing combined assurance: Insights from multiple case studies
L. Decaux and G. Sarens · 2014
Earlier work this paper cites.
Risk assessment and risk management: Review of recent advances on their foundation
T. Aven · 2015
Earlier work this paper cites.
Combined assurance: One language, one voice, one view
S. C. Huibers · 2015
Earlier work this paper cites.
Accounting for unknown unknowns in managing multi-hazard risks
R. B. Gilbert, M. Habibi, and F. Nadim · 2016
Earlier work this paper cites.
Defense-in-depth
J.-E. Holmberg · 2017
Cited alongside, same era.
Three Lines of Defence: A robust organising framework, or just lines in the sand?
H. Davies and M. Zhivitskaya · 2018
Cited alongside, same era.
Risk management — Guidelines
ISO 31000 · 2018
Cited alongside, same era.
Defence in depth against human extinction: Prevention, response, resilience, and why they all matter
O. Cotton-Barratt, M. Daniel, and A. Sandberg · 2020
Cited alongside, same era.
The IIA’s Three Lines Model: An update of the Three Lines of Defense
Institute of Internal Auditors · 2020
Cited alongside, same era.
Good and bad reasons: The Swiss cheese model and its critics
J. Larouzee and J.-C. Le Coze · 2020
Cited alongside, same era.
Model evaluation for extreme risks
T. Shevlane, S. Farquhar, B. Garfinkel, M. Phuong, J. Whittlestone, J. Leung, D. Kokotajlo, N. Marchal, M. Anderljung, N. Kolt, L. Ho, D. Siddarth, S. Avin, W. Hawkins, B. Kim, I. Gabriel, V. Bolina, J. Clark, Y. Bengio, P. Christiano, and A. Dafoe · 2023
Later among the works it cites.
Responsible scaling: Comparing government guidance and company policy
B. Anderson-Samways, S. Ee, J. O’Brien, M. Buhl, and Z. Williams · 2024
Closest in time.
Reflections on our Responsible Scaling Policy
Anthropic · 2024
Closest in time.
Introducing the Frontier Safety Framework
A. Dragan, K. King, Helen, and A. Dafoe · 2024
Closest in time.
Frontier AI Safety Commitments, AI Seoul Summit 2024
DSIT · 2024
Closest in time.
Adapting cybersecurity frameworks to manage frontier AI risks: A defense-in-depth approach
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Coordination challenges in implementing the Three Lines of Defense model
U. Bantleon, A. d’Arcy, M. Eulerich, A. Hucke, B. Pedell, and N. V. Ratzinger-Sakel · 2021
Cited alongside, same era.
AI certification: Advancing ethical practice by reducing information asymmetries
P. Cihon, M. J. Kleinaltenkamp, J. Schuett, and S. D. Baum · 2021
Cited alongside, same era.
Coordinated pausing: An evaluation-based coordination scheme for frontier AI developers
J. Alaga and J. Schuett · 2023
Cited alongside, same era.
Responsible Scaling Policy
Anthropic · 2023
Cited alongside, same era.
Emerging Processes for Frontier AI Safety
DSIT · 2023
Cited alongside, same era.
Information technology — Artificial intelligence — Guidance on risk management
ISO/IEC 23894 · 2023
Cited alongside, same era.
S. Ee, J. O’Brien, Z. Williams, A. El-Dakhakhni, M. Aird, and A. Lintz · 2024
Closest in time.
Risk thresholds for frontier AI
L. Koessler, J. Schuett, and M. Anderljung · 2024
Closest in time.
Algorithmic black swans
N. Kolt · 2024
Closest in time.
Responsible reporting for frontier AI development
N. Kolt, M. Anderljung, J. Barnhart, A. Brass, K. Esvelt, G. K. Hadfield, L. Heim, M. Rodriguez, J. B. Sandbrink, and T. Woodside · 2024
Closest in time.
AGI Readiness Policy Version 1.0
Magic · 2024
Closest in time.
Common elements of frontier AI safety policies
METR · 2024
Closest in time.
Do companies’ AI safety policies meet government best practice?
S. Ó hÉigeartaigh, Y. Lannquist, A. Marcoci, J. Sevilla, M. A. Ulloa Ruiz, Y. Chaudhary, T. Schreier, Z. Stein-Perlman, and J. L. Ladish · 2024
Closest in time.
Evaluating frontier models for dangerous capabilities
M. Phuong, M. Aitchison, E. Catt, S. Cogan, A. Kaskasoli, V. Krakovna, D. Lindner, M. Rahtz, Y. Assael, S. Hodkinson, H. Howard, T. Lieberum, R. Kumar, M. A. Raad, A. Webson, L. Ho, S. Lin, S. Farquhar, M. Hutter, G. Deletang, A. Ruoss, S. El-Sayed, S. Brown, A. Dragan, R. Shah, A. Dafoe, and T. Shevlane · 2024
Closest in time.
Transforming risk governance at frontier AI companies
B. Robinson and J. Ginns · 2024
Closest in time.
Is OpenAI’s Preparedness Framework better than its competitors’ Responsible Scaling Policies? A comparative analysis
SaferAI · 2024
Closest in time.
From principles to rules: A regulatory approach for frontier AI
J. Schuett, M. Anderljung, A. Carlier, L. Koessler, and B. Garfinkel · 2024
Closest in time.
Scaling AI safely: Can preparedness frameworks pull their weight?
J. Titus · 2024
Closest in time.