Fetching the paper…
Reading the bibliography…
Frontier artificial intelligence (AI) systems pose increasing risks to society, making it essential for developers to provide assurances about their safety.
The goal structuring notation: A safety argument notation
T. Kelly and R. Weaver · 2004
Earlier work this paper cites.
Reviewing assurance arguments: A step-by-step approach
T. Kelly · 2007
Earlier work this paper cites.
Has the safety case failed?
B. Fitzgerald, P. Breen, and J. Patrick · 2010
Earlier work this paper cites.
Software certification: Is there a case against safety cases?
A. Wassyng, T. Maibaum, M. Lawford, and H. Bherer · 2011
Earlier work this paper cites.
Building blocks for assurance cases
R. Bloomfield and K. Netkachova · 2014
Earlier work this paper cites.
Safety cases
T. Kelly · 2017
Earlier work this paper cites.
Modelling confidence in railway safety case
R. Wang, J. Guiochet, G. Motet, and W. Schön · 2017
Earlier work this paper cites.
The Agile Safety Case
T. Myklebust and T. Stålhane · 2018
Earlier work this paper cites.
CAP 670: Air traffic services safety requirements
CAA · 2019
Earlier work this paper cites.
The role of safety architectures in aviation safety cases
E. Denney, G. Pai, and I. Whiteside · 2019
Earlier work this paper cites.
Implementation of nuclear safety cases
A. Bounds · 2020
Earlier work this paper cites.
Autonomous cars, trust and safety case for the public
T. Myklebust, T. Stålhane, G. D. Jenssen, and I. Wærø · 2020
Earlier work this paper cites.
Deepfakes and disinformation: Exploring the impact of synthetic political video on deception, uncertainty, and trust in news
C. Vaccari and A. Chadwick · 2020
Earlier work this paper cites.
Safety case templates for autonomous systems
R. Bloomfield, G. Fletcher, H. Khlaaf, L. Hinde, and P. Ryan · 2021
Earlier work this paper cites.
M. R. Rahman, R. Mahdavi-Hezaveh, and L. Williams · 2021
Earlier work this paper cites.
Goal Structuring Notation Community Standard Version 3
SCSC · 2021
Earlier work this paper cites.
The emerging threat of AI-driven cyber attacks: A review
B. Guembe, A. Azeta, S. Misra, V. C. Osamor, L. Fernandez-Sanz, and V. Pospelova · 2022
Earlier work this paper cites.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre · 2022
Earlier work this paper cites.
Langchain
LangChain · 2022
Earlier work this paper cites.
Protecting society from AI misuse: When are restrictions on capabilities warranted?
M. Anderljung and J. Hazell · 2023
Earlier work this paper cites.
Scheming AIs: Will AIs fake alignment during training in order to get power?
J. Carlsmith · 2023
Earlier work this paper cites.
Harms from increasingly agentic algorithmic systems
A. Chan, R. Salganik, A. Markelius, C. Pang, N. Rajkumar, D. Krasheninnikov, L. Langosco, Z. He, Y. Duan, M. Carroll, M. Lin, A. Mayhew, K. Collins, M. Molamohammadi, J. Burden, W. Zhao, S. Rismani, K. Voudouris, U. Bhatt, A. Weller, D. Krueger, and T. Maharaj · 2023
Earlier work this paper cites.
AI capabilities can be significantly improved without expensive retraining
T. Davidson, J.-S. Denain, P. Villalobos, and G. Bas · 2023
Earlier work this paper cites.
Spear phishing with large language models
J. Hazell · 2023
Earlier work this paper cites.
Toward comprehensive risk assessments and assurance of AI-based systems
H. Khlaaf · 2023
Earlier work this paper cites.
AgentBench: Evaluating LLMs as agents
X. Liu, H. Yu, H. Zhang, Y. Xu, X. Lei, H. Lai, Y. Gu, H. Ding, K. Men, K. Yang, S. Zhang, X. Deng, A. Zeng, Z. Du, C. Zhang, S. Shen, T. Zhang, Y. Su, H. Sun, M. Huang, Y. Dong, and J. Tang · 2023
Earlier work this paper cites.
GAIA: a benchmark for general AI asssistants
G. Mialon, C. Fourrier, C. Swift, T. Wolf, Y. LeCun, and T. Scialom · 2023
Earlier work this paper cites.
Auditing large language models: A three-layered approach
J. Mökander, J. Schuett, H. R. Kirk, and L. Floridi · 2023
Cited alongside, same era.
AI deception: A survey of examples, risks, and potential solutions
P. S. Park, S. Goldstein, A. O’Gara, M. Chen, and D. Hendrycks · 2023
Cited alongside, same era.
Are emergent abilities of large language models a mirage?
R. Schaeffer, B. Miranda, and S. Koyejo · 2023
Cited alongside, same era.
Reflexion: Language agents with verbal reinforcement learning
N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao · 2023
Cited alongside, same era.
Can large language models democratize access to dual-use biotechnology?
E. H. Soice, R. Rocha, K. Cordova, M. Specter, and K. M. Esvelt · 2023
Sleeper agents: Training deceptive LLMs that persist through safety training
E. Hubinger, C. Denison, J. Mu, M. Lambert, M. Tong, M. MacDiarmid, T. Lanham, D. M. Ziegler, T. Maxwell, N. Cheng, A. Jermyn, A. Askell, A. Radhakrishnan, C. Anil, D. Duvenaud, D. Ganguli, F. Barez, J. Clark, K. Ndousse, K. Sachan, M. Sellitto, M. Sharma, N. DasSarma, R. Grosse, S. Kravec, Y. Bai, Z. Witten, M. Favaro, J. Brauner, H. Karnofsky, P. Christiano, S. R. Bowman, L. Graham, J. Kaplan, S. Mindermann, R. Greenblatt, B. Shlegeris, N. Schiefer, and E. Perez · 2024
Closest in time.
Safety cases at AISI
G. Irving · 2024
Closest in time.
SWE-bench: Can language models resolve real-world GitHub issues?
C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan · 2024
Closest in time.
Assuring AI safety: fallible knowledge and the gricean maxims
M. H. L. Kaas and I. Habli · 2024
Closest in time.
Evaluating language-model agents on realistic autonomous tasks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Cyber threat intelligence mining for proactive cybersecurity defense: A survey and new perspectives
N. Sun, M. Ding, J. Jiang, W. Xu, X. Mo, Y. Tai, and J. Zhang · 2023
Cited alongside, same era.
Tree of thoughts: Deliberate problem solving with large language models
S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan · 2023
Cited alongside, same era.
A new initiative for developing third-party model evaluations
Anthropic · 2024
Cited alongside, same era.
Three sketches of ASL-4 safety case components
Anthropic · 2024
Cited alongside, same era.
Towards evaluations-based safety cases for AI scheming
M. Balesni, M. Hobbhahn, D. Lindner, A. Meinke, T. Korbak, J. Clymer, B. Shlegeris, J. Scheurer, C. Stix, R. Shah, N. Goldowsky-Dill, D. Braun, B. Chughtai, O. Evans, D. Kokotajlo, and L. Bushnaq · 2024
Cited alongside, same era.
International scientific report on the safety of advanced AI
Y. Bengio, P. Daniel, B. Tamay, B. Rishi, C. Stephen, C. Yejin, G. Danielle, H. Hoda, K. Leila, L. Shayne, V. Mavroudis, M. Mazeika, K. Y. Ng, C. T. Okolo, D. Raji, T. Skeadas, and F. Tramèr · 2024
Cited alongside, same era.
Managing extreme AI risks amid rapid progress
Y. Bengio, G. Hinton, A. Yao, D. Song, P. Abbeel, T. Darrell, Y. N. Harari, Y.-Q. Zhang, L. Xue, S. Shalev-Shwartz, G. Hadfield, J. Clune, T. Maharaj, F. Hutter, A. G. Baydin, S. McIlraith, Q. Gao, A. Acharya, D. Krueger, A. Dragan, P. Torr, S. Russell, D. Kahneman, J. Brauner, and S. Mindermann · 2024
Cited alongside, same era.
M. Kinniment, L. J. K. Sato, H. Du, B. Goodrich, M. Hasin, L. Chan, L. H. Miles, T. R. Lin, H. Wijk, J. Burget, A. Ho, E. Barnes, and P. Christiano · 2024
Closest in time.
Risk thresholds for frontier AI
L. Koessler, J. Schuett, and M. Anderljung · 2024
Closest in time.
Guidelines for capability elicitation
METR · 2024
Closest in time.
Portable evaluation tasks via the METR task standard
METR · 2024
Closest in time.
Resources
METR · 2024
Closest in time.
Critical national infrastructure (CNI)
NCSC · 2024
Closest in time.
The near-term impact of AI on the cyber threat
NCSC · 2024
Closest in time.
Securing ai model weights: Preventing theft and misuse of frontier models
S. Nevo, D. Lahav, A. Karpur, Y. Bar-On, H. A. Bradley, and J. Alstott · 2024
Closest in time.
The alignment problem from a deep learning perspective
R. Ngo, L. Chan, and S. Mindermann · 2024
Closest in time.
Cyber attack glossary
NIST · 2024
Closest in time.
OpenAI o1 system card
OpenAI · 2024
Closest in time.
Reconciling kaplan and chinchilla scaling laws
T. Pearce and J. Song · 2024
Closest in time.
Evaluating frontier models for dangerous capabilities
M. Phuong, M. Aitchison, E. Catt, S. Cogan, A. Kaskasoli, V. Krakovna, D. Lindner, M. Rahtz, Y. Assael, S. Hodkinson, H. Howard, T. Lieberum, R. Kumar, M. A. Raad, A. Webson, L. Ho, S. Lin, S. Farquhar, M. Hutter, G. Deletang, A. Ruoss, S. El-Sayed, S. Brown, A. Dragan, R. Shah, A. Dafoe, and T. Shevlane · 2024
Closest in time.
Beyond chinchilla-optimal: Accounting for inference in language model scaling laws
N. Sardana, J. Portes, S. Doubov, and J. Frankle · 2024
Closest in time.
From principles to rules: A regulatory approach for frontier ai
J. Schuett, M. Anderljung, A. Carlier, L. Koessler, and B. Garfinkel · 2024
Closest in time.
2023 state of deepfakes: Realities, threats, and impact
Security Hero · 2024
Closest in time.
AI sandbagging: Language models can strategically underperform on evaluations
T. van der Weij, F. Hofstätter, O. Jaffe, S. F. Brown, and F. R. Ward · 2024
Closest in time.
Affirmative safety: An approach to risk management for advanced AI
A. Wasil, J. Clymer, D. Krueger, E. Dardaman, S. Campos, and E. Murphy · 2024
Closest in time.
SWE-agent: Agent-computer interfaces enable automated software engineering
J. Yang, C. E. Jimenez, A. Wettig, K. Lieret, S. Yao, K. Narasimhan, and O. Press · 2024
Closest in time.
A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly
Y. Yao, J. Duan, K. Xu, Y. Cai, Z. Sun, and Y. Zhang · 2024
Closest in time.
Cybench: A framework for evaluating cybersecurity capabilities and risks of language models
A. K. Zhang, N. Perry, R. Dulepet, J. Ji, J. W. Lin, E. Jones, C. Menders, G. Hussein, S. Liu, D. Jasper, P. Peetathawatchai, A. Glenn, V. Sivashankar, D. Zamoshchin, L. Glikbarg, D. Askaryar, M. Yang, T. Zhang, R. Alluri, N. Tran, R. Sangpisit, P. Yiorkadjis, K. Osele, G. Raghupathi, D. Boneh, D. E. Ho, and P. Liang · 2024
Closest in time.
When LLMs meet cybersecurity: A systematic literature review
J. Zhang, H. Bu, H. Wen, Y. Chen, L. Li, and H. Zhu · 2024
Closest in time.