Fetching the paper…
Reading the bibliography…
Organizations that develop and deploy artificial intelligence (AI) systems need to take measures to reduce the associated risks.
Effective whistle-blowing
J. P. Near and M. P. Miceli · 1995
Earlier work this paper cites.
Whistleblowing: A restrictive definition and interpretation
P. B. Jubb · 1999
Earlier work this paper cites.
Existential risks: Analyzing human extinction scenarios and related hazards
N. Bostrom · 2001
Earlier work this paper cites.
The Black Swan: The Impact of the Highly Improbable
N. N. Taleb · 2007
Earlier work this paper cites.
Cognitive biases potentially affecting judgment of global risks
E. Yudkowsky · 2008
Earlier work this paper cites.
What makes whistleblowing effective: Whistleblowing in Peru and South Korea
C. R. Apaza and Y. Chang · 2011
Earlier work this paper cites.
Information hazards: A typology of potential harms from knowledge
N. Bostrom · 2011
Earlier work this paper cites.
Risk governance
M. B. Van Asselt and O. Renn · 2011
Earlier work this paper cites.
On the meaning of a black swan in a risk context
T. Aven · 2013
Earlier work this paper cites.
Workplace bullying after whistleblowing: Future research and implications
B. Bjørkelo · 2013
Earlier work this paper cites.
Guide 51:2014 Safety aspects — Guidelines for their inclusion in standards, 2014
ISO/IEC · 2014
Earlier work this paper cites.
Quadratic voting as efficient corporate governance
E. A. Posner and E. G. Weyl · 2014
Earlier work this paper cites.
The psychology of whistleblowing
J. Dungan, A. Waytz, and L. Young · 2015
Earlier work this paper cites.
Why firms implement risk governance: Stepping beyond traditional risk management to enterprise risk management
S. A. Lundqvist · 2015
Earlier work this paper cites.
Concrete problems in AI safety
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Earlier work this paper cites.
Racing to the precipice: A model of artificial intelligence development
S. Armstrong, N. Bostrom, and C. Shulman · 2016
Earlier work this paper cites.
Driving priorities in risk-based regulation: What’s the problem?
R. Baldwin and J. Black · 2016
Earlier work this paper cites.
The malicious use of artificial intelligence: Forecasting, prevention, and mitigation
M. Brundage, S. Avin, J. Clark, H. Toner, P. Eckersley, B. Garfinkel, A. Dafoe, P. Scharre, T. Zeitzoff, B. Filar, H. Anderson, H. Roff, G. C. Allen, J. Steinhardt, C. Flynn, S. O. hÉigeartaigh, S. Beard, H. Belfield, S. Farquhar, C. Lyle, R. Crootof, O. Evans, M. Page, J. Bryson, R. Yampolskiy, and D. Amodei · 2018
Earlier work this paper cites.
An AI race for strategic advantage: Rhetoric and risks
S. Cave and S. S. ÓhÉigeartaigh · 2018
Earlier work this paper cites.
Google is helping the Pentagon build AI for drones
K. Conger and D. Cameron · 2018
Earlier work this paper cites.
Three lines of defence: A robust organising framework, or just lines in the sand?
H. Davies and M. Zhivitskaya · 2018
Earlier work this paper cites.
31000:2018 Risk management — Guidelines, 2018
ISO · 2018
Earlier work this paper cites.
Quadratic voting: How mechanism design can radicalize democracy
S. P. Lalley and E. G. Weyl · 2018
Earlier work this paper cites.
https://twitter.com/ssnstudy/status/1112099054551515138 , 2019
A. Aquisti · 2019
Earlier work this paper cites.
First report of the Axon AI & Policing Technology Ethics Board, 2019
Axon · 2019
Earlier work this paper cites.
The vulnerable world hypothesis
N. Bostrom · 2019
Earlier work this paper cites.
Artificial intelligence research needs responsible publication norms
R. Crootof · 2019
Earlier work this paper cites.
Googlers against transphobia and hate
Googlers Against Transphobia · 2019
Earlier work this paper cites.
Anchors aweigh: The sources, variety, and challenges of mission drift
M. G. Grimes, T. A. Williams, and E. Y. Zhao · 2019
Earlier work this paper cites.
31010:2019 Risk management — Risk assessment techniques, 2019
IEC · 2019
Earlier work this paper cites.
The global landscape of AI ethics guidelines
A. Jobin, M. Ienca, and E. Vayena · 2019
Earlier work this paper cites.
Designing artificial intelligence review boards: creating risk metrics for review of AI
S. R. Jordan · 2019
Earlier work this paper cites.
How viable is international arms control for military artificial intelligence? three lessons from nuclear weapons
M. M. Maas · 2019
Earlier work this paper cites.
Principles alone cannot guarantee ethical AI
B. Mittelstadt · 2019
Earlier work this paper cites.
Google’s brand-new AI ethics board is already falling apart
K. Piper · 2019
Earlier work this paper cites.
Actionable auditing: Investigating the impact of publicly naming biased performance results of commercial AI products
I. D. Raji and J. Buolamwini · 2019
Earlier work this paper cites.
Building data and AI ethics committees
R. Sandler, J. Basl, and S. Tiell · 2019
Earlier work this paper cites.
Release strategies and the social impacts of language models
I. Solaiman, M. Brundage, J. Clark, A. Askell, A. Herbert-Voss, J. Wu, A. Radford, G. Krueger, J. W. Kim, S. Kreps, M. McCain, A. Newhouse, J. Blazakis, K. McGuffie, and J. Wang · 2019
Earlier work this paper cites.
Independence with a purpose: Facebook’s creative use of Delaware’s purpose trust statute to establish independent oversight
V. Thomas, J. Duda, and T. Maurer · 2019
Earlier work this paper cites.
Create an ethics committee to keep your AI initiative in check
S. Tiell · 2019
Earlier work this paper cites.
An external advisory council to help advance the responsible development of AI
K. Walker · 2019
Cited alongside, same era.
Activism by the AI community: Analysing recent achievements and future prospects
H. Belfield · 2020
Cited alongside, same era.
From ethics washing to ethics bashing: A view on tech ethics from within moral philosophy
E. Bietti · 2020
Cited alongside, same era.
Toward trustworthy AI development: Mechanisms for supporting verifiable claims
M. Brundage, S. Avin, J. Wang, H. Belfield, G. Krueger, G. Hadfield, H. Khlaaf, J. Yang, H. Toner, R. Fong, T. Maharaj, P. W. Koh, S. Hooker, J. Leung, A. Trask, E. Bluemke, J. Lebensold, C. O’Keefe, M. Koren, T. Ryffel, J. Rubinovitz, T. Besiroglu, F. Carugati, J. Clark, P. Eckersley, S. de Haas, M. Johnson, B. Laurie, A. Ingerman, I. Krawczuk, A. Askell, R. Cammarota, A. Lohn, D. Krueger, C. Stix, P. Henderson, L. Graham, C. Prunkl, B. Martin, E. Seger, N. Zilberman, S. O. hÉigeartaigh, F. Kroeger, G. Sastry, R. Kagan, A. Weller, B. Tse, E. Barnes, A. Dafoe, P. Scharre, A. Herbert-Voss, M. Rasser, S. Sodhani, C. Flynn, T. K. Gilbert, L. Dyer, S. Khan, Y. Bengio, and M. Anderljung · 2020
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
D. Ganguli, L. Lovitt, J. Kernion, A. Askell, Y. Bai, S. Kadavath, B. Mann, E. Perez, N. Schiefer, K. Ndousse, A. Jones, S. Bowman, A. Chen, T. Conerly, N. DasSarma, D. Drain, N. Elhage, S. El-Showk, S. Fort, Z. Hatfield-Dodds, T. Henighan, D. Hernandez, T. Hume, J. Jacobson, S. Johnston, S. Kravec, C. Olsson, S. Ringer, E. Tran-Johnson, D. Amodei, T. Brown, N. Joseph, S. McCandlish, C. Olah, J. Kaplan, and J. Clark · 2022
Later among the works it cites.
Unsolved problems in ML safety
D. Hendrycks, N. Carlini, J. Schulman, and J. Steinhardt · 2022
Later among the works it cites.
How our principles helped define alphafold’s release
K. Kavukcuoglu, P. Kohli, L. Ibrahim, D. Bloxwich, and S. Brown · 2022
Later among the works it cites.
Conjecture: Internal infohazard policy
C. Leahy, S. Black, C. Scammell, and A. Miotti · 2022
Later among the works it cites.
Operationalising AI governance through ethics-based auditing: An industry case study
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Defence in depth against human extinction: Prevention, response, resilience, and why they all matter
O. Cotton-Barratt, M. Daniel, and A. Sandberg · 2020
Cited alongside, same era.
Negotiating ’evil’: Google, Project Maven and the corporate form
P. Crofts and H. van Rijswijk · 2020
Cited alongside, same era.
AI ethics groups are repeating one of society’s classic mistakes
A. Gupta and V. Heath · 2020
Cited alongside, same era.
The ethics of AI ethics: An evaluation of guidelines
T. Hagendorff · 2020
Cited alongside, same era.
The flight to safety-critical AI
W. Hunt · 2020
Cited alongside, same era.
The Facebook Oversight Board: Creating an independent institution to adjudicate online free expression
K. Klonick · 2020
Cited alongside, same era.
Putting principles into practice: How we approach responsible AI at Microsoft
Microsoft · 2020
Cited alongside, same era.
Decolonial AI: Decolonial theory as sociotechnical foresight in artificial intelligence
S. Mohamed, M.-T. Png, and W. Isaac · 2020
Cited alongside, same era.
J. Mökander and L. Floridi · 2022
Later among the works it cites.
Best practices for deploying language models
OpenAI · 2022
Later among the works it cites.
Securing ongoing funding
Oversight Board · 2022
Later among the works it cites.
Red teaming language models with language models
E. Perez, S. Huang, F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving · 2022
Later among the works it cites.
Looking before we leap
M. Petermann, N. Tempini, I. K. Garcia, K. Whitaker, and A. Strait · 2022
Later among the works it cites.
Outsider oversight: Designing a third party audit ecosystem for AI governance
I. D. Raji, P. Xu, C. Honigsberg, and D. Ho · 2022
Later among the works it cites.
Red-teaming the Stable Diffusion safety filter
J. Rando, D. Paleka, D. Lindner, L. Heim, and F. Tramèr · 2022
Later among the works it cites.
Differential technology development: A responsible innovation principle for navigating technology risks
J. Sandbrink, H. Hobbs, J. Swett, A. Dafoe, and A. Sandberg · 2022
Later among the works it cites.
Three lines of defense against risks from AI
J. Schuett · 2022
Later among the works it cites.
In defence of principlism in AI ethics and governance
E. Seger · 2022
Later among the works it cites.
Compute trends across three eras of machine learning
J. Sevilla, L. Heim, A. Ho, T. Besiroglu, M. Hobbhahn, and P. Villalobos · 2022
Later among the works it cites.
Structured access: An emerging paradigm for safe AI deployment
T. Shevlane · 2022
Later among the works it cites.
Axon committed to listening and learning so that we can fulfill our mission to protect life, together
R. Smith · 2022
Later among the works it cites.
Advancing ethics review practices in AI research
M. Srikumar, R. Finlay, G. Abuhamad, C. Ashurst, R. Campbell, E. Campbell-Ratcliffe, H. Hongo, S. R. Jordan, J. Lindley, A. Ovadya, et al · 2022
Later among the works it cites.
Dual use of artificial-intelligence-powered drug discovery
F. Urbina, F. Lentzos, C. Invernizzi, and S. Ekins · 2022
Later among the works it cites.
Meta’s Oversight Board: A review and critical assessment
D. Wong and L. Floridi · 2022
Later among the works it cites.
AI ethics: From principles to practice
J. Zhou and F. Chen · 2022
Later among the works it cites.
Protecting society from AI misuse: When are restrictions on capabilities warranted?
M. Anderljung and J. Hazell · 2023
Closest in time.
Core views on AI safety: When, why, what, and how
Anthropic · 2023
Closest in time.
J. A. Goldstein, G. Sastry, M. Musser, R. DiResta, M. Gentzel, and K. Sedova · 2023
Closest in time.
Google calls in help from Larry Page and Sergey Brin for A.I. fight
N. Grant · 2023
Closest in time.
Microsoft eyes $10 billion bet on ChatGPT
L. Hoffman and R. Albergotti · 2023
Closest in time.
23894:2023 Information technology — Artificial intelligence — Guidance on risk management, 2023
ISO/IEC · 2023
Closest in time.
Algorithmic black swans
N. Kolt · 2023
Closest in time.
Microsoft and OpenAI extend partnership
Microsoft · 2023
Closest in time.
Our approach
Microsoft · 2023
Closest in time.
Auditing large language models: A three-layered approach
J. Mökander, J. Schuett, H. R. Kirk, and L. Floridi · 2023
Closest in time.
The alignment problem from a deep learning perspective
R. Ngo, L. Chan, and S. Mindermann · 2023
Closest in time.
Artificial Intelligence Risk Management Framework (AI RMF 1.0), 2023
NIST · 2023
Closest in time.
OpenAI · 2023
Closest in time.
OpenAI and Microsoft extend partnership, 2023
OpenAI · 2023
Closest in time.
https://www.oversightboard.com , 2023
Oversight Board · 2023
Closest in time.
Our commitment
Oversight Board · 2023
Closest in time.
Trustees
Oversight Board · 2023
Closest in time.
Risk management in the Artificial Intelligence Act
J. Schuett · 2023
Closest in time.
The gradient of generative AI release: Methods and considerations
I. Solaiman · 2023
Closest in time.