Fetching the paper…
Reading the bibliography…
We present a quantitative model for tracking dangerous AI capabilities over time.
“The Role of Cooperation in Responsible AI Development”
Amanda Askell, Miles Brundage and Gillian Hadfield · 1907
Earlier work this paper cites.
“The Role of Cooperation in Responsible AI Development”
Amanda Askell, Miles Brundage and Gillian Hadfield · 1907
Earlier work this paper cites.
“Racing to the Precipice: A Model of Artificial Intelligence Development”
Stuart Armstrong, Nick Bostrom and Carl Shulman · 2016
Earlier work this paper cites.
“Racing to the Precipice: A Model of Artificial Intelligence Development”
Stuart Armstrong, Nick Bostrom and Carl Shulman · 2016
Earlier work this paper cites.
“To Regulate or Not: A Social Dynamics Analysis of an Idealised AI Race”
The Han, Luis Pereira, Francisco. Santos and Tom Lenaerts · 2020
Earlier work this paper cites.
“Specification Gaming: The Flip Side of AI Ingenuity” Retrieved February 2023 from https://deepmind.com/blog/article/Specification-gaming-the-flip-side-of-AI-ingenuity, 2020
Victoria Krakovna et al · 2020
Earlier work this paper cites.
“To Regulate or Not: A Social Dynamics Analysis of an Idealised AI Race”
The Han, Luis Pereira, Francisco. Santos and Tom Lenaerts · 2020
Earlier work this paper cites.
“Specification Gaming: The Flip Side of AI Ingenuity” Retrieved February 2023 from https://deepmind.com/blog/article/Specification-gaming-the-flip-side-of-AI-ingenuity, 2020
Victoria Krakovna et al · 2020
Earlier work this paper cites.
“Why and How Governments Should Monitor AI Development”
Jess Whittlestone and Jack Clark · 2021
Earlier work this paper cites.
“Why and How Governments Should Monitor AI Development”
Jess Whittlestone and Jack Clark · 2021
Earlier work this paper cites.
“Social diversity reduces the complexity and cost of fostering fairness”
Theodor Cimpeanu, Alessandro Di Stefano, Cedric Perret and The Han · 2022
Earlier work this paper cites.
“Voluntary Safety Commitments Provide an Escape from Over-Regulation in AI Development”
The Han, Tom Lenaerts, Francisco Santos and Luís Pereira · 2022
Earlier work this paper cites.
“Unsolved Problems in ML Safety”, 2022
Dan Hendrycks, Nicholas Carlini, John Schulman and Jacob Steinhardt · 2022
Earlier work this paper cites.
“Social diversity reduces the complexity and cost of fostering fairness”
Theodor Cimpeanu, Alessandro Di Stefano, Cedric Perret and The Han · 2022
Earlier work this paper cites.
“Voluntary Safety Commitments Provide an Escape from Over-Regulation in AI Development”
The Han, Tom Lenaerts, Francisco Santos and Luís Pereira · 2022
Earlier work this paper cites.
“Unsolved Problems in ML Safety”, 2022
Dan Hendrycks, Nicholas Carlini, John Schulman and Jacob Steinhardt · 2022
Earlier work this paper cites.
“Coordinated pausing: An evaluation-based coordination scheme for frontier AI developers”, 2023
Jide Alaga and Jonas Schuett · 2023
Earlier work this paper cites.
“Who is leading in AI? An analysis of industry AI research”, 2023
Ben Cottier, Tamay Besiroglu and David Owen · 2023
Earlier work this paper cites.
“An International Consortium for Evaluations of Societal-Scale Risks from Advanced AI”
Ross Gruetzemacher et al · 2023
Earlier work this paper cites.
“An Overview of Catastrophic AI Risks”
Dan Hendrycks, Mantas Mazeika and Thomas Woodside · 2023
Earlier work this paper cites.
“International Institutions for Advanced AI”
Lewis Ho et al · 2023
Earlier work this paper cites.
“Industrial Policy for Advanced AI: Compute Pricing and the Safety Tax”
Mckay Jensen, Nicholas Emery-Xu and Robert Trager · 2023
Earlier work this paper cites.
“Responsible Scaling Policies (RSPs)”
METR · 2023
Earlier work this paper cites.
“Please Report Your Compute”
Jaime Sevilla · 2023
Earlier work this paper cites.
“Model Evaluation for Extreme Risks”
Toby Shevlane et al · 2023
Earlier work this paper cites.
“Coordinated pausing: An evaluation-based coordination scheme for frontier AI developers”, 2023
Jide Alaga and Jonas Schuett · 2023
Earlier work this paper cites.
“Who is leading in AI? An analysis of industry AI research”, 2023
Ben Cottier, Tamay Besiroglu and David Owen · 2023
Cited alongside, same era.
“An International Consortium for Evaluations of Societal-Scale Risks from Advanced AI”
Ross Gruetzemacher et al · 2023
Cited alongside, same era.
“An Overview of Catastrophic AI Risks”
Dan Hendrycks, Mantas Mazeika and Thomas Woodside · 2023
Cited alongside, same era.
“International Institutions for Advanced AI”
Lewis Ho et al · 2023
Cited alongside, same era.
“Industrial Policy for Advanced AI: Compute Pricing and the Safety Tax”
Mckay Jensen, Nicholas Emery-Xu and Robert Trager · 2023
Cited alongside, same era.
“Responsible Scaling Policies (RSPs)”
METR · 2023
Cited alongside, same era.
“Open Problems in Technical AI Governance”
Anka Reuel et al · 2024
Closest in time.
“Pre-Deployment Evaluation of Anthropic’s Upgraded Claude 3.5 Sonnet”, 2024
UK AI Safety Institute · 2024
Closest in time.
“Pre-Deployment Evaluation of Open AI’s o1 model”, 2024
UK AI Safety Institute · 2024
Closest in time.
“Managing Extreme AI Risks amid Rapid Progress”
Yoshua Bengio et al · 2024
Closest in time.
“International Scientific Report on the Safety of Advanced AI (Interim Report)”
Yoshua Bengio et al · 2024
Closest in time.
“Sabotage Evaluations for Frontier Models”
Joe Benton et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Please Report Your Compute”
Jaime Sevilla · 2023
Cited alongside, same era.
“Model Evaluation for Extreme Risks”
Toby Shevlane et al · 2023
Cited alongside, same era.
“Managing Extreme AI Risks amid Rapid Progress”
Yoshua Bengio et al · 2024
Cited alongside, same era.
“International Scientific Report on the Safety of Advanced AI (Interim Report)”
Yoshua Bengio et al · 2024
Cited alongside, same era.
“Sabotage Evaluations for Frontier Models”
Joe Benton et al · 2024
Cited alongside, same era.
“Both eyes open: Vigilant Incentives help auditors improve AI safety”
Paolo Bova, Alessandro Di and The Han · 2024
Cited alongside, same era.
Paolo Bova, Alessandro Di and The Han · 2024
Closest in time.
“Safety Cases for Frontier AI”
Marie Buhl et al · 2024
Closest in time.
“Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems”
David"davidad" Dalrymple et al · 2024
Closest in time.
“Safety Case Template for Frontier AI: A Cyber Inability Argument”
Arthur Goemans et al · 2024
Closest in time.
“Analyzing Probabilistic Methods for Evaluating Agent Capabilities”
Axel Højmark et al · 2024
Closest in time.
“Evaluating Language-Model Agents on Realistic Autonomous Tasks”
Megan Kinniment et al · 2024
Closest in time.
“Risk Thresholds for Frontier AI”
Leonie Koessler, Jonas Schuett and Markus Anderljung · 2024
Closest in time.
“Details about METR’s Preliminary Evaluation of OpenAI O1-Preview”
METR · 2024
Closest in time.
“Example Protocol”
METR · 2024
Closest in time.
“Secret Collusion among Generative AI Agents”, 2024
Sumeet Motwani et al · 2024
Closest in time.
“FACT SHEET: U.S. Department of Commerce & U.S. Department of State Launch the International Network of AI Safety Institutes at Inaugural Convening in San Francisco”
NIST · 2024
Closest in time.
“OpenAI o1 System Card”, 2024
OpenAI · 2024
Closest in time.
“AI Deception: A Survey of Examples, Risks, and Potential Solutions”
Peter. Park et al · 2024
Closest in time.
“Evaluating Frontier Models for Dangerous Capabilities”
Mary Phuong et al · 2024
Closest in time.
“Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?”
Richard Ren et al · 2024
Closest in time.
“Open Problems in Technical AI Governance”
Anka Reuel et al · 2024
Closest in time.
“Pre-Deployment Evaluation of Anthropic’s Upgraded Claude 3.5 Sonnet”, 2024
UK AI Safety Institute · 2024
Closest in time.
“Pre-Deployment Evaluation of Open AI’s o1 model”, 2024
UK AI Safety Institute · 2024
Closest in time.
“Can AI Scaling Continue Through 2030?” Accessed: 2024-12-18, 2024
Jaime Sevilla et al · 2030
Closest in time.
“Can AI Scaling Continue Through 2030?” Accessed: 2024-12-18, 2024
Jaime Sevilla et al · 2030
Closest in time.