Fetching the paper…
Reading the bibliography…
Existing strategies for managing risks from advanced AI systems often focus on affecting what AI systems are developed and how they diffuse.
Annex B: Glossary of Terms
International Panel on Climate Change. 2001 · 2001
Earlier work this paper cites.
Remedying Election Wrongs
Huefner, S. 2007 · 2007
Earlier work this paper cites.
Automation Bias in Intelligent Time Critical Decision Support Systems
Cummings, M. L. 2012 · 2012
Earlier work this paper cites.
The Making of Arduino
Kushner, D. 2013 · 2013
Earlier work this paper cites.
Damage in the English Law of Negligence
Nolan, D. 2013 · 2013
Earlier work this paper cites.
Annex II: Glossary
International Panel on Climate Change. 2014 · 2014
Earlier work this paper cites.
AI Governance: A Research Agenda
Dafoe, A. 2018 · 2018
Earlier work this paper cites.
Private Accountability in the Age of Artificial Intelligence
Katyal, S. K. 2018 · 2018
Earlier work this paper cites.
Understanding Deterrence
Mazarr, M. J. 2018 · 2018
Earlier work this paper cites.
Automation and New Tasks: How Technology Displaces and Reinstates Labor
Acemoglu, D.; and Restrepo, P. 2019 · 2019
Earlier work this paper cites.
Deep Fakes: A Looming Challenge for Privacy, Democracy, and National Security
Chesney, R.; and Citron, D. K. 2019 · 2019
Earlier work this paper cites.
Deep Neural Network Fingerprinting by Conferrable Adversarial Examples
Lukas; Zhang; and Kerschbaum. 2019 · 2019
Earlier work this paper cites.
Model Cards for Model Reporting
Mitchell, M.; Wu, S.; Zaldivar, A.; Barnes, P.; Vasserman, L.; Hutchinson, B.; Spitzer, E.; Raji, I. D.; and Gebru, T. 2019 · 2019
Earlier work this paper cites.
What is AI Literacy? Competencies and Design Considerations
Long, D.; and Magerko, B. 2020 · 2020
Earlier work this paper cites.
Countering the Cyber Enforcement Gap: Strengthening Global Capacity on Cybercrime
Peters, A.; and Jordan, A. 2020 · 2020
Earlier work this paper cites.
Datasheets for Datasets
Gebru, T.; Morgenstern, J.; Vecchione, B.; Vaughan, J. W.; Wallach, H.; Daumé III, H.; and Crawford, K. 2021 · 2021
Earlier work this paper cites.
Monitoring AI Services for Misuse
Javadi, S. A.; Norval, C.; Cloete, R.; and Singh, J. 2021 · 2021
Earlier work this paper cites.
AI and Shared Prosperity
Klinova, K.; and Korinek, A. 2021 · 2021
Earlier work this paper cites.
Preventing Repeated Real World AI Failures by Cataloging Incidents: The AI Incident Database
McGregor, S. 2021 · 2021
Earlier work this paper cites.
Governing Algorithmic Systems with Impact Assessments: Six Observations
Watkins, E. A.; Moss, E.; Metcalf, J.; Singh, R.; and Elish, M. C. 2021 · 2021
Earlier work this paper cites.
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Bai, Y.; Jones, A.; Ndousse, K.; Askell, A.; Chen, A.; DasSarma, N.; Drain, D.; Fort, S.; Ganguli, D.; Henighan, T.; Joseph, N.; et al. 2022 · 2022
Earlier work this paper cites.
Measuring Progress on Scalable Oversight for Large Language Models
Bowman, S.; Hyun, J.; Perez, E.; Chen, E.; Pettit, C.; Heiner, S.; Lukošiūtė, K.; Askell, A.; Jones, A.; Chen, A.; et al. 2022 · 2022
Earlier work this paper cites.
Is Power-Seeking AI an Existential Risk?
Carlsmith, J. 2022 · 2022
Earlier work this paper cites.
A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values
Clarke, S.; and Whittlestone, J. 2022 · 2022
Earlier work this paper cites.
Developing and implementing 20-mph speed limits in Edinburgh and Belfast: mixed-methods study
Jepson, R.; Baker, G.; Cleland, C.; Cope, A.; Craig, N.; Foster, C.; Hunter, R.; Kee, F.; Kelly, M. P.; Kelly, P.; Milton, K.; Nightingale, G.; Turner, K.; Williams, A. J.; and Woodcock, K. 2022 · 2022
Earlier work this paper cites.
Human-Centred Mechanism Design with Democratic AI
Koster, R.; Balaguer, J.; Tacchetti, A.; Weinstein, A.; Zhu, T.; Hauser, O.; Williams, D.; Campbell-Gillingham, L.; Thacker, P.; Botvinick, M.; and Summerfield, C. 2022 · 2022
Earlier work this paper cites.
How Open Source Machine Learning Software Shapes AI
Langenkamp, M.; and Yue, D. N. 2022 · 2022
Earlier work this paper cites.
Will AI Make Cyber Swords or Shields?
Lohn, A. J.; and Jackson, K. A. 2022 · 2022
Earlier work this paper cites.
Outsider Oversight: Designing a Third Party Audit Ecosystem for AI Governance
Raji, I. D.; Xu, P.; Honigsberg, C.; and Ho, D. 2022 · 2022
Earlier work this paper cites.
Structured Access: An Emerging Paradigm for Safe AI Deployment
Shevlane, T. 2022 · 2022
Earlier work this paper cites.
Birdwatch: Crowd Wisdom and Bridging Algorithms Can Inform Understanding and Reduce the Spread of Misinformation
Wojcik, S.; Hilgard, S.; Judd, N.; Mocanu, D.; Ragain, S.; Hunzaker, M. B.; Coleman, K.; and Baxter, J. 2022 · 2022
Earlier work this paper cites.
Protecting Society from AI Misuse: When are Restrictions on Capabilities Warranted?
Anderljung, M.; and Hazell, J. 2023 · 2023
Earlier work this paper cites.
Will Humanity Choose Its Future?
Assadi, G. 2023 · 2023
Earlier work this paper cites.
Structured Access for Third-Party Research on Frontier AI Models: Investigating Researchers’ Model Access Requirements
Bucknall, B. S.; and Trager, R. F. 2023 · 2023
Earlier work this paper cites.
Fighting Misinformation with Authenticated C2PA Provenance Metadata
Earnshaw, D.; and MacCormack. 2023 · 2023
Earlier work this paper cites.
Adapting Cybersecurity Frameworks to Manage Frontier AI Risks: a Defense-in-Depth Approach
Ee, S. 2023 · 2023
Earlier work this paper cites.
The Stable Signature: Rooting Watermarks in Latent Diffusion Models
Fernandez, P.; Couairon, G.; Jégou, H.; Douze, M.; and Furon, T. 2023 · 2023
Earlier work this paper cites.
BadLlama: Cheaply Removing Safety Fine-Tuning from Llama 2-Chat 13B
Gade, P.; Lermen, S.; Rogers-Smith, C.; and Ladish, J. 2023 · 2023
Earlier work this paper cites.
An Overview of Catastrophic AI Risks
Hendrycks; Mazeika; and Woodside. 2023 · 2023
Earlier work this paper cites.
Natural Selection Favors AIs over Humans
Hendrycks, D. 2023 · 2023
Earlier work this paper cites.
International Institutions for Advanced AI
Ho, L.; Barnhart, J.; Trager, R.; Bengio, Y.; Brundage, M.; Carnegie, A.; Chowdhury, R.; Dafoe, A.; Hadfield, G.; Levi, M.; and Snidal, D. 2023 · 2023
Earlier work this paper cites.
AI safety on whose terms?
Lazar, S.; and Nelson, A. 2023 · 2023
Earlier work this paper cites.
Report Launch: Examining Risks at the Intersection of AI and Bio
Nelson, C.; and Rose, S. 2023 · 2023
Cited alongside, same era.
Securing Artificial Intelligence Model Weights
Nevo, S.; Lahav, D.; Karpur, A.; Alstott, J.; and Matheny, J. 2023 · 2023
Cited alongside, same era.
GPT-4 Technical Report: System Card
OpenAI. 2023 · 2023
Cited alongside, same era.
Increased Compute Efficiency and the Diffusion of AI Capabilities
Pilz, K.; Heim, L.; and Brown, N. 2023 · 2023
Cited alongside, same era.
The Liar’s Dividend: Can Politicians Claim Misinformation to Evade Accountability?
Schiff; Schiff; and Bueno. 2023 · 2023
Cited alongside, same era.
Open Sourcing Highly Capable Foundation Models
Seger, E.; Dreksler, N.; Moulange, R.; Dardaman, E.; Schuett, J.; Wei, K.; Winter, C.; Arnold, M.; Ó hÉigeartaigh, S.; Korinek, A.; Anderljung, M.; Bucknall, B.; Chan, A.; Stafford, E.; Koessler, L.; Ovadya, A.; Garfinkel, B.; Bluemke, E.; Aird, M.; Levermore, P.; Hazell, J.; and Gupta, A. 2023 · 2023
On the Societal Impact of Open Foundation Models
Kapoor, S.; Bommasani, R.; Klyman, K.; Longpre, S.; Ramaswami, A.; Cihon, P.; Hopkins, A.; Bankston, K.; Biderman, S.; Bogen, M.; Chowdhury, R.; Engler, A.; Henderson, P.; Jernite, Y.; Lazar, S.; Maffulli, S.; Pineau, J.; Skowron, A.; Song, D.; Storchan, V.; Ho, D. E.; Liang, P.; and Narayanan, A. 2024 · 2024
Closest in time.
Malawi’s Re-Run Election is Lesson for African Opposition
Kell, F. 2020 · 2024
Closest in time.
Responsible Reporting for Frontier AI Development
Kolt, N.; Mazeika, M.; Barnhart, J.; Brass, A.; Esvelt, K.; Hadfield, G. K.; Heim, L.; Rodriguez, M.; Sandbrink, J. B.; and Woodside, T. 2024 · 2024
Closest in time.
Models on the Frontline: AI’s Defensive Role
Krier, S. 2024 · 2024
Closest in time.
A Revealing Picture
Lakatos, S. 2023 · 2024
Closest in time.
Top State Officials Push to Make Spread of US Election Misinformation Illegal
Lerner, K. 2023 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Please Report Your Compute
Sevilla, J.; Ho, A.; and Besiroglu, T. 2023 · 2023
Cited alongside, same era.
Model Evaluation for Extreme Risks
Shevlane, T.; Farquhar, S.; Garfinkel, B.; Phuong, M.; Whittlestone, J.; Leung, J.; Kokotajlo, D.; Marchal, N.; Anderljung, M.; Kolt, N.; Ho, L.; Siddarth, D.; Avin, S.; Hawkins, W.; Kim, B.; Gabriel, I.; Bolina, V.; Clark, J.; Bengio, Y.; Christiano, P.; and Dafoe, A. 2023 · 2023
Cited alongside, same era.
The Gradient of Generative AI Release: Methods and Considerations
Solaiman, I. 2023 · 2023
Cited alongside, same era.
Skating to Where the Puck Is Going: Anticipating and Managing Risks from Frontier AI Systems
Toner, H.; Ji, J.; Bansemer, J.; Lim, L.; Painter, C. D.; Corley, C. D.; Whittlestone, J.; Botvinick, M.; Rodriguez, M.; and Kumar, R. S. S. 2023 · 2023
Cited alongside, same era.
Market Concentration Implications of Foundation Models: The Invisible Hand of ChatGPT
Vipra, J.; and Korinek, A. 2023 · 2023
Cited alongside, same era.
Sociotechnical Safety Evaluation of Generative AI Systems
Weidinger, L.; Rauh, M.; Marchal, N.; Manzini, A.; Hendricks, L. A.; Mateos-Garcia, J.; Bergman, S.; Kay, J.; Griffin, C.; Bariach, B.; Gabriel, I.; Rieser, V.; and Isaac, W. 2023 · 2023
Cited alongside, same era.
Closest in time.
Demystifying GPT-3
Li, C. 2020 · 2024
Closest in time.
The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
Li, N.; Pan, A.; Gopal, A.; Yue, S.; Berrios, D.; Gatti, A.; Li, J. D.; Dombrowski, A.-K.; Goel, S.; Phan, L.; Mukobi, G.; Helm-Burger, N.; Lababidi, R.; Justen, L.; Liu, A. B.; Chen, M.; Barrass, I.; Zhang, O.; Zhu, X.; Tamirisa, R.; Bharathi, B.; Khoja, A.; Zhao, Z.; Herbert-Voss, A.; Breuer, C. B.; Zou, A.; Mazeika, M.; Oswal, P.; Liu, W.; Hunt, A. A.; Tienken-Harder, J.; Shih, K. Y.; Talley, K.; Guan, J.; Kaplan, R.; Steneker, I.; Campbell, D.; Jokubaitis, B.; Levinson, A.; Wang, J.; Qian, W.; Karmakar, K. K.; Basart, S.; Fitz, S.; Levine, M.; Kumaraguru, P.; Tupakula, U.; Varadharajan, V.; Shoshitaishvili, Y.; Ba, J.; Esvelt, K. M.; Wang, A.; and Hendrycks, D. 2024 · 2024
Closest in time.
Rishi Sunak Promised to Make AI Safe. Big Tech’s Not Playing Ball
Manancourt; Volpicelli; and Chatterjee. 2024 · 2024
Closest in time.
Berlin Stages Partial Rerun of 2021 German Federal Election
Martin, N.; Hallam, M.; and Hubenko, D. 2024 · 2024
Closest in time.
Turkey’s Deep Fake-Influenced Election Spells Trouble
Meyer, D. 2023 · 2024
Closest in time.
Microsoft and OpenAI Launch Societal Resilience Fund
Microsoft. 2024 · 2024
Closest in time.
Cybersecurity and AI: The Evolving Security Landscape
Newman, S. 2024 · 2024
Closest in time.
OpenAI Charter
OpenAI. 2018 · 2024
Closest in time.
Climate Finance and the USD 100 Billion Goal
Organisation for Economic Co-operation and Development (OECD). 2023 · 2024
Closest in time.
Artificial Intelligence on the Corporate Board?
Pugh, W. 2019 · 2024
Closest in time.
Tracking Compute-Intensive AI Models
Rahman; Owen; ; and You. 2024 · 2024
Closest in time.
Escalation Risks from Language Models in Military and Diplomatic Decision-Making
Rivera, J. P.; Mukobi, G.; Reuel, A.; Lamparth, M.; Smith, C.; and Schneider, J. 2024 · 2024
Closest in time.
No, LLM Agents Can Not Autonomously Exploit One-Day Vulnerabilities
Rohlf, C. 2024 · 2024
Closest in time.
Meet My A.I. Friends
Roose, K. 2024 · 2024
Closest in time.
On the Conversational Persuasiveness of Large Language Models: A Randomized Controlled Trial
Salvi, F.; Horta Ribeiro, M.; Gallotti, R.; and West, R. 2024 · 2024
Closest in time.
A Quarter of Europeans Want AI to Replace Politicians. That’s a Terrible Idea
Samuel, S. 2019 · 2024
Closest in time.
Computing Power and the Governance of Artificial Intelligence
Sastry, G.; Heim, L.; Belfield, H.; Anderljung, M.; Brundage, M.; Hazell, J.; O’Keefe, C.; Hadfield, G. K.; Ngo, R.; Pilz, K.; Gor, G.; Bluemke, E.; Shoker, S.; Egan, J.; Trager, R. F.; Avin, S.; Weller, A.; Bengio, Y.; and Coyle, D. 2024 · 2024
Closest in time.
Future-Proofing Frontier AI Regulation: Projecting Future Compute for Frontier AI Models
Scharre, P. 2024 · 2024
Closest in time.
Why Proof of Humanity Is More Important Than Ever
Shoemaker, P. 2024 · 2024
Closest in time.
Detecting AI Fingerprints: A Guide to Watermarking and Beyond
Srinivasan, S. 2024 · 2024
Closest in time.
The Role of Governments in Increasing Interconnected Post-Deployment Monitoring of AI
Stein, M.; Bernardi, J.; and Dunlop, C. 2024 · 2024
Closest in time.
New Hampshire Investigating Fake Biden Robocall Meant to Discourage Voters Ahead of Primary
Swenson, A.; and Weissert, W. 2024 · 2024
Closest in time.
Press Statement: ACLU of Georgia Opposes Bill Criminalizing ’Deep Fakes’ About Election Candidates
Toney, D. 2024 · 2024
Closest in time.
AI Foundation Models Initial Report
U.K. Competition and Markets Authority. 2023 · 2024
Closest in time.
Emerging Processes for Frontier AI Safety
U.K. Department for Science, Innovation and Technology. 2023 · 2024
Closest in time.
Reported Road Casualties Great Britain: Pedestrian Factsheet 2021
U.K. Department for Transport. 2022 · 2024
Closest in time.
FDA’s Drug Review Process: Ensuring Drugs Are Safe and Effective
U.S. Food and Drug Administration. 2017 · 2024
Closest in time.
Block Nuclear Launch by Autonomous Artificial Intelligence Act of 2023
U.S. Senate. 2023 · 2024
Closest in time.
GPT-3 Quality for 500k
Venigalla, A.; and Li, L. 2022 · 2024
Closest in time.
Meta’s Powerful AI Language Model Has Leaked Online — What Happens Now?
Vincent, J. 2023 · 2024
Closest in time.
Merging AI Incidents Research with Political Misinformation Research: Introducing the Political Deepfakes Incidents Database
Walker; Schiff; and Schiff. 2024 · 2024
Closest in time.
Managing Risks - Cause, Event and Consequence
Waycott, A. 2018 · 2024
Closest in time.
How AI Will Transform the 2024 Elections
West, D. M. 2023 · 2024
Closest in time.
Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence
White House. 2023 · 2024
Closest in time.
WILDCHAT: 1M ChatGPT Interaction Logs in the Wild
Zhao, W.; Ren, X.; Hessel, J.; Cardie, C.; Choi, Y.; and Deng, Y. 2024 · 2024
Closest in time.
The Hidden Risk of Letting AI Decide – Losing the Skills to Choose for Ourselves
Árvai, J. 2024 · 2024
Closest in time.