Fetching the paper…
Reading the bibliography…
Mitigating the risks from frontier AI systems requires up-to-date and reliable information about those systems.
“The Role of Cooperation in Responsible AI Development”
Amanda Askell, Miles Brundage and Gillian Hadfield · 1907
Earlier work this paper cites.
“Bottlenecks and baselines: Tackling information deficits in environmental regulation”
Bradley Karkkainen · 2008
Earlier work this paper cites.
“Information Acquisition and Institutional Design”
Matthew. Stephenson · 2011
Earlier work this paper cites.
“Airline Safety Improvement Through Experience with Near-Misses: A Cautionary Tale”
Peter Madsen, Robin. Dillon and Catherine. Tinsley · 2015
Earlier work this paper cites.
“Model Cards for Model Reporting”
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Raji and Timnit Gebru · 2019
Earlier work this paper cites.
“Regulatory Monitors: Policing Firms in the Compliance Era”
Rory Van · 2019
Earlier work this paper cites.
“Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims”
Miles Brundage, Shahar Avin, Jasmine Wang, Haydn Belfield, Gretchen Krueger, Gillian Hadfield, Heidy Khlaaf, Jingying Yang, Helen Toner, Ruth Fong, Tegan Maharaj, Pang Koh, Sara Hooker, Jade Leung, Andrew Trask, Emma Bluemke, Jonathan Lebensold, Cullen O’Keefe, Mark Koren, Théo Ryffel, J.. Rubinovitz, Tamay Besiroglu, Federica Carugati, Jack Clark, Peter Eckersley, Sarah de Haas, Maritza Johnson, Ben Laurie, Alex Ingerman, Igor Krawczuk, Amanda Askell, Rosario Cammarota, Andrew Lohn, David Krueger, Charlotte Stix, Peter Henderson, Logan Graham, Carina Prunkl, Bianca Martin, Elizabeth Seger, Noa Zilberman, Seán hÉigeartaigh, Frens Kroeger, Girish Sastry, Rebecca Kagan, Adrian Weller, Brian Tse, Elizabeth Barnes, Allan Dafoe, Paul Scharre, Ariel Herbert-Voss, Martijn Rasser, Shagun Sodhani, Carrick Flynn, Thomas Gilbert, Lisa Dyer, Saif Khan, Yoshua Bengio and Markus Anderljung · 2020
Earlier work this paper cites.
“AI Accidents: An Emerging Threat”, 2021
Zachary Arnold and Helen Toner · 2021
Earlier work this paper cites.
“Filling gaps in trustworthy development of AI” Publisher: American Association for the Advancement of Science
Shahar Avin, Haydn Belfield, Miles Brundage, Gretchen Krueger, Jasmine Wang, Adrian Weller, Markus Anderljung, Igor Krawczuk, David Krueger, Jonathan Lebensold, Tegan Maharaj and Noa Zilberman · 2021
Earlier work this paper cites.
“2021/0106(COD): Artificial Intelligence Act”
European Parliament · 2021
Earlier work this paper cites.
“Governing AI safety through independent audits” Publisher: Nature Publishing Group
Gregory Falco, Ben Shneiderman, Julia Badger, Ryan Carrier, Anton Dahbura, David Danks, Martin Eling, Alwyn Goodloe, Jerry Gupta, Christopher Hart, Marina Jirotka, Henric Johnson, Cara LaPointe, Ashley. Llorens, Alan. Mackworth, Carsten Maple, Sigurður Pálsson, Frank Pasquale, Alan Winfield and Zee Yeong · 2021
Earlier work this paper cites.
“Datasheets for Datasets”
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Vaughan, Hanna Wallach, Hal DauméIII and Kate Crawford · 2021
Earlier work this paper cites.
“Preventing Repeated Real World AI Failures by Cataloging Incidents: The AI Incident Database” Number: 17
Sean McGregor · 2021
Earlier work this paper cites.
“AI Risk Management Framework”
National Institute of Standards and Technology · 2021
Earlier work this paper cites.
“Why and How Governments Should Monitor AI Development”
Jess Whittlestone and Jack Clark · 2021
Earlier work this paper cites.
“CFPB Supervision and Examination Process”, 2022
CFPB · 2022
Earlier work this paper cites.
“System Safety and Artificial Intelligence”
Roel.. Dobbe · 2022
Earlier work this paper cites.
“Predictability and Surprise in Large Generative Models”
Deep Ganguli, Danny Hernandez, Liane Lovitt, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova Dassarma, Dawn Drain, Nelson Elhage, Sheer El, Stanislav Fort, Zac Hatfield-Dodds, Tom Henighan, Scott Johnston, Andy Jones, Nicholas Joseph, Jackson Kernian, Shauna Kravec, Ben Mann, Neel Nanda, Kamal Ndousse, Catherine Olsson, Daniela Amodei, Tom Brown, Jared Kaplan, Sam McCandlish, Christopher Olah, Dario Amodei and Jack Clark · 2022
Earlier work this paper cites.
“Unsolved Problems in ML Safety”
Dan Hendrycks, Nicholas Carlini, John Schulman and Jacob Steinhardt · 2022
Earlier work this paper cites.
“Soft Law 2.0: An Agile and Effective Governance Approach for Artificial Intelligence”
Gary. Marchant and Carlos Gutierrez · 2022
Earlier work this paper cites.
“Indexing AI Risks with Incidents, Issues, and Variants”
Sean McGregor, Kevin Paeth and Khoa Lam · 2022
Earlier work this paper cites.
“Disclosure by Design: Designing information disclosures to support meaningful transparency and accountability”
Chris Norval, Kristin Cornelius, Jennifer Cobbe and Jatinder Singh · 2022
Earlier work this paper cites.
“Outsider Oversight: Designing a Third Party Audit Ecosystem for AI Governance”
Inioluwa Raji, Peggy Xu, Colleen Honigsberg and Daniel Ho · 2022
Earlier work this paper cites.
“Blueprint for an AI Bill of Rights | OSTP”, 2022
The White House · 2022
Earlier work this paper cites.
“Emergent Abilities of Large Language Models”
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean and William Fedus · 2022
Earlier work this paper cites.
“Comment of the AI Policy and Governance Working Group on the NTIA AI Accountability Policy Request for Comment Docket NTIA-230407-0093”, 2023
AI Policy and Governance Working Group · 2023
Earlier work this paper cites.
“Frontier AI Regulation: Managing Emerging Risks to Public Safety”
Markus Anderljung, Joslyn Barnhart, Anton Korinek, Jade Leung, Cullen O’Keefe, Jess Whittlestone, Shahar Avin, Miles Brundage, Justin Bullock, Duncan Cass-Beggs, Ben Chang, Tantum Collins, Tim Fist, Gillian Hadfield, Alan Hayes, Lewis Ho, Sara Hooker, Eric Horvitz, Noam Kolt, Jonas Schuett, Yonadav Shavit, Divya Siddarth, Robert Trager and Kevin Wolf · 2023
Earlier work this paper cites.
“Towards Publicly Accountable Frontier LLMs: Building an External Scrutiny Ecosystem under the ASPIRE Framework”
Markus Anderljung, Everett Smith, Joe O’Brien, Lisa Soder, Benjamin Bucknall, Emma Bluemke, Jonas Schuett, Robert Trager, Lacey Strahm and Rumman Chowdhury · 2023
Earlier work this paper cites.
“Anthropic’s Responsible Scaling Policy”, 2023
Anthropic · 2023
Cited alongside, same era.
“Managing AI Risks in an Era of Rapid Progress”
Yoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Yuval Harari, Ya-Qin Zhang, Lan Xue, Shai Shalev-Shwartz, Gillian Hadfield, Jeff Clune, Tegan Maharaj, Frank Hutter, Atılımüneş Baydin, Sheila McIlraith, Qiqi Gao, Ashwin Acharya, David Krueger, Anca Dragan, Philip Torr, Stuart Russell, Daniel Kahneman, Jan Brauner and Sören Mindermann · 2023
Cited alongside, same era.
“The Foundation Model Transparency Index”
Rishi Bommasani, Kevin Klyman, Shayne Longpre, Sayash Kapoor, Nestor Maslej, Betty Xiong, Daniel Zhang and Percy Liang · 2023
Cited alongside, same era.
“Ecosystem Graphs: The Social Footprint of Foundation Models”
Rishi Bommasani, Dilara Soylu, Thomas. Liao, Kathleen. Creel and Percy Liang · 2023
Cited alongside, same era.
“Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback”
Stephen Casper, Xander Davies, Claudia Shi, Thomas Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, Tony Wang, Samuel Marks, Charbel-Raphaël Segerie, Micah Carroll, Andi Peng, Phillip Christoffersen, Mehul Damani, Stewart Slocum, Usman Anwar, Anand Siththaranjan, Max Nadeau, Eric. Michaud, Jacob Pfau, Dmitrii Krasheninnikov, Xin Chen, Lauro Langosco, Peter Hase, Erdem Bıyık, Anca Dragan, David Krueger, Dorsa Sadigh and Dylan Hadfield-Menell · 2023
“Towards best practices in AGI safety and governance: A survey of expert opinion”
Jonas Schuett, Noemi Dreksler, Markus Anderljung, David McCaffary, Lennart Heim, Emma Bluemke and Ben Garfinkel · 2023
Later among the works it cites.
“Policy paper: Introducing the AI Safety Institute” ISBN: 978-1-5286-4538-6, 2023
Secretary of State for Science, Innovation and Technology · 2023
Later among the works it cites.
“Model evaluation for extreme risks”
Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, Lewis Ho, Divya Siddarth, Shahar Avin, Will Hawkins, Been Kim, Iason Gabriel, Vijay Bolina, Jack Clark, Yoshua Bengio, Paul Christiano and Allan Dafoe · 2023
Later among the works it cites.
“How to deal with an AI near-miss: Look to the skies”
Kris Shrishak · 2023
Later among the works it cites.
“Biden-Harris Administration Launches Artificial Intelligence Cyber Challenge to Protect America’s Critical Software”, 2023
The White House · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Information Markets and AI Development”
Jack Clark · 2023
Cited alongside, same era.
“AI Governance: Overview and Theoretical Lenses”
Allan Dafoe · 2023
Cited alongside, same era.
“Policy paper: Emerging processes for frontier AI safety”, 2023
Department for Science, Innovation and Technology · 2023
Cited alongside, same era.
“Policy Updates”, 2023
Department for Science, Innovation and Technology and AI Safety Institute · 2023
Cited alongside, same era.
“Tech entrepreneur Ian Hogarth to lead UK’s AI Foundation Model Taskforce”, 2023
Department for Science, Innovation and Technology, AI Safety Institute, Chloe Smith MP and The Rt Hon Rishi Sunak MP · 2023
Cited alongside, same era.
“The Bletchley Declaration by Countries Attending the AI Safety Summit, 1-2 November 2023”, 2023
Department for Science, Innovation and Technology, Foreign, Commonwealth \& Development Office and Prime Minister’s Office, 10 Downing Street · 2023
Cited alongside, same era.
“Oversight for Frontier AI through a Know-Your-Customer Scheme for Compute Providers”
Janet Egan and Lennart Heim · 2023
Cited alongside, same era.
“Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence”, 2023
The White House · 2023
Later among the works it cites.
“FACT SHEET: Biden-Harris Administration Announces New Actions to Promote Responsible AI Innovation that Protects Americans’ Rights and Safety”, 2023
The White House · 2023
Later among the works it cites.
“FACT SHEET: Biden-Harris Administration Secures Voluntary Commitments from Leading Artificial Intelligence Companies to Manage the Risks Posed by AI”, 2023
The White House · 2023
Later among the works it cites.
“Skating to Where the Puck Is Going”, 2023
Helen Toner, Jessica Ji, John Bansemer and Lucy Lim · 2023
Later among the works it cites.
“Why We Need to Know More: Exploring the State of AI Incident Documentation Practices”
Violet Turri and Rachel Dzombak · 2023
Later among the works it cites.
“At the Direction of President Biden, Department of Commerce to Establish U.S. Artificial Intelligence Safety Institute to Lead Efforts on AI Safety”, 2023
U.S. Department of Commerce · 2023
Later among the works it cites.
“Sociotechnical Safety Evaluation of Generative AI Systems”
Laura Weidinger, Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Hendricks, Juan Mateos-Garcia, Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, Iason Gabriel, Verena Rieser and William Isaac · 2023
Later among the works it cites.
“Responsible Scaling: Comparing Government Guidance and Company Policy”, 2024
Bill Anderson-Samways, Shaun Ee, Joe O’Brien, Marie Buhl and Zoe Williams · 2024
Closest in time.
“AI auditing: The Broken Bus on the Road to AI Accountability”
Abeba Birhane, Ryan Steed, Victor Ojewale, Briana Vecchione and Inioluwa Raji · 2024
Closest in time.
“Black-Box Access is Insufficient for Rigorous AI Audits”
Stephen Casper, Carson Ezell, Charlotte Siegmann, Noam Kolt, Taylor Curtis, Benjamin Bucknall, Andreas Haupt, Kevin Wei, Jérémy Scheurer, Marius Hobbhahn, Lee Sharkey, Satyapriya Krishna, Marvin Von, Silas Alberti, Alan Chan, Qinyi Sun, Michael Gerovitch, David Bau, Max Tegmark, David Krueger and Dylan Hadfield-Menell · 2024
Closest in time.
“Epoch Database”, 2024
EPOCH · 2024
Closest in time.
“AI Regulation Has Its Own Alignment Problem: The Technical and Institutional Feasibility of Disclosure, Registration, Licensing, and Auditing”
Neel Guha, Christie Lawrence, Lindsey Gailmard, Kit Rodolfa, Faiz Surani, Rishi Bommasani, Inioluwa Raji, Mariano-Florentino Cuéllar, Colleen Honigsberg, Percy Liang and Daniel Ho · 2024
Closest in time.
“Governing Through the Cloud: The Intermediary Role of Compute Providers in AI Regulation”
Lennart Heim, Tim Fist, Janet Egan, Sihao Huang, Stephen Zekany, Robert Trager, Michael. Osborne and Noa Zilberman · 2024
Closest in time.
“On the Societal Impact of Open Foundation Models”, 2024
Sayash Kapoor, Rishi Bommasani, Kevin Klyman, Shayne Longpre, Ashwin Ramaswami, Peter Cihon, Aspen Hopkins, Kevin Bankston, Stella Biderman, Miranda Bogen, Rumman Chowdhury, Alex Engler, Peter Henderson, Yacine Jernite, Seth Lazar, Stefano Maffulli, Alondra Nelson, Joelle Pineau, Aviya Skowron, Dawn Song, Victor Storchan, Daniel Zhang, Daniel Ho, Percy Liang and Arvind Narayanan · 2024
Closest in time.
“Governing AI Agents”, 2024
Noam Kolt · 2024
Closest in time.
“A Safe Harbor for AI Evaluation and Red Teaming”
Shayne Longpre, Sayash Kapoor, Kevin Klyman, Ashwin Ramaswami, Rishi Bommasani, Borhane Blili-Hamelin, Yangsibo Huang, Aviya Skowron, Zheng-Xin Yong, Suhas Kotha, Yi Zeng, Weiyan Shi, Xianjun Yang, Reid Southen, Alexander Robey, Patrick Chao, Diyi Yang, Ruoxi Jia, Daniel Kang, Sandy Pentland, Arvind Narayanan, Percy Liang and Peter Henderson · 2024
Closest in time.
“Towards AI Accountability Infrastructure: Gaps and Opportunities in AI Audit Tooling”
Victor Ojewale, Ryan Steed, Briana Vecchione, Abeba Birhane and Inioluwa Raji · 2024
Closest in time.
“Evaluating Frontier Models for Dangerous Capabilities”
Mary Phuong, Matthew Aitchison, Elliot Catt, Sarah Cogan, Alexandre Kaskasoli, Victoria Krakovna, David Lindner, Matthew Rahtz, Yannis Assael, Sarah Hodkinson, Heidi Howard, Tom Lieberum, Ramana Kumar, Maria Raad, Albert Webson, Lewis Ho, Sharon Lin, Sebastian Farquhar, Marcus Hutter, Gregoire Deletang, Anian Ruoss, Seliem El-Sayed, Sasha Brown, Anca Dragan, Rohin Shah, Allan Dafoe and Toby Shevlane · 2024
Closest in time.
“Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context”
Machel Reid et al · 2024
Closest in time.
“Computing Power and the Governance of Artificial Intelligence”
Girish Sastry, Lennart Heim, Haydn Belfield, Markus Anderljung, Miles Brundage, Julian Hazell, Cullen O’Keefe, Gillian. Hadfield, Richard Ngo, Konstantin Pilz, George Gor, Emma Bluemke, Sarah Shoker, Janet Egan, Robert. Trager, Shahar Avin, Adrian Weller, Yoshua Bengio and Diane Coyle · 2024
Closest in time.
“Biden-Harris Administration Announces First-Ever Consortium Dedicated to AI Safety”, 2024
U.S. Department of Commerce · 2024
Closest in time.
“Keeping Up with the Frontier: Why Congress Should Codify Reporting Requirements For Advanced AI Systems”, 2024
Thomas Woodside · 2024
Closest in time.