Fetching the paper…
Reading the bibliography…
Artificial Intelligence (AI) is progressing rapidly, and companies are shifting their focus to developing generalist AI systems that can autonomously act and pursue goals.
“Social Choice Theory”
Amartya Sen · 1986
Earlier work this paper cites.
“The Science of Computing: The Internet Worm”
Peter Denning · 1989
Earlier work this paper cites.
“Deep Blue”
Murray Campbell, A Hoane and Feng-Hsiung Hsu · 2002
Earlier work this paper cites.
“A Systematic Approach to Safety Case Management”
Tim Kelly · 2004
Earlier work this paper cites.
“A Comparative Review of Risk Management Standards”
Tzvi Raz and David Hillson · 2005
Earlier work this paper cites.
“The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization”, 2020
Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt and Justin Gilmer · 2006
Earlier work this paper cites.
“EAD Safety Case Guidance”, 2010
European Organisation for the Safety of Air Navigation · 2010
Earlier work this paper cites.
“Infusion Pumps Total Product Life Cycle - Guidance for Industry and FDA Staff”, 2014
Food and Drug Administration · 2014
Earlier work this paper cites.
“The Off-Switch Game”
Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel and Stuart Russell · 2017
Earlier work this paper cites.
“Automating Inequality: How High-Tech Tools Profile, Police and Punish the Poor”
Virginia Eubanks · 2018
Earlier work this paper cites.
“Superhuman AI for multiplayer poker”
Noam Brown and Tuomas Sandholm · 2019
Earlier work this paper cites.
“Incomplete contracting and AI alignment”
Dylan Hadfield-Menell and Gillian Hadfield · 2019
Earlier work this paper cites.
“Optimal policies tend to seek power”
A Turner, L Smith, R Shah and A Critch · 2019
Earlier work this paper cites.
“Toward Agile Governance: The Pattern of Emerging Industry Development and Regulation”
Lan Xue and Jing Zhao · 2019
Earlier work this paper cites.
“Consequences of misaligned AI”
Simon Zhuang and Dylan Hadfield-Menell · 2020
Earlier work this paper cites.
“Safety of artificial intelligence: A collaborative model”, 2020
J Mcdermid and Yan Jia · 2020
Earlier work this paper cites.
“A graph placement methodology for fast chip design”
Azalia Mirhoseini, Anna Goldie, Mustafa Yazgan, Joe Jiang, Ebrahim Songhori, Shen Wang, Young-Joon Lee, Eric Johnson, Omkar Pathak, Azade Nazi, Jiwoo Pak, Andy Tong, Kavya Srinivasa, William Hang, Emre Tuncer, Quoc Le, James Laudon, Richard Ho, Roger Carpenter and Jeff Dean · 2021
Earlier work this paper cites.
“Highly accurate protein structure prediction with AlphaFold”
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, Alex Bridgland, Clemens Meyer, Simon Kohl, Andrew Ballard, Andrew Cowie, Bernardino Romera-Paredes, Stanislav Nikolov, Rishub Jain, Jonas Adler, Trevor Back, Stig Petersen, David Reiman, Ellen Clancy, Michal Zielinski, Martin Steinegger, Michalina Pacholska, Tamas Berghammer, Sebastian Bodenstein, David Silver, Oriol Vinyals, Andrew Senior, Koray Kavukcuoglu, Pushmeet Kohli and Demis Hassabis · 2021
Earlier work this paper cites.
“On the Opportunities and Risks of Foundation Models”, 2021
Rishi Bommasani et al · 2021
Earlier work this paper cites.
“Unsolved Problems in ML Safety”, 2021
Dan Hendrycks, Nicholas Carlini, John Schulman and Jacob Steinhardt · 2021
Earlier work this paper cites.
“Using Advance Market Commitments for Public Purpose Technology Development”, Policy Brief, 2021
Alan Ho and Jake Taylor · 2021
Earlier work this paper cites.
“Alphabet annual report, page 33 (page 71 in the pdf): ‘As of December 31, 2022, we had USD113.8 billion in cash, cash equivalents, and short-term marketable securities’. [For comparison, the cost of training GPT-4 has been estimated as USD50 million (https://epochai.org/trends), and Sam Altman, the CEO of OpenAI, has stated that the cost for the whole process was more than USD100 million (https://www.wired.com/story/openai-ceo-sam-altman-the-age-of-giant-ai-models-is-already-over/).]”, https://abc.xyz/assets/d4/4f/a48b94d548d0b2fdc029a95e8c63/2022-alphabet-annual-report.pdf , 2022
Alphabet · 2022
Earlier work this paper cites.
“Algorithmic progress in computer vision”, 2022
Ege Erdil and Tamay Besiroglu · 2022
Earlier work this paper cites.
“ML-Enhanced Code Completion Improves Developer Productivity” Accessed: 2023-9-15, https://blog.research.google/2022/07/ml-enhanced-code-completion-improves.html
Maxim Tabachnyk · 2022
Earlier work this paper cites.
“Constitutional AI: Harmlessness from AI Feedback”, 2022
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Kamile Lukosuite, Liane Lovitt, Michael Sellitto, Nelson Elhage, Nicholas Schiefer, Noemi Mercado, Nova DasSarma, Robert Lasenby, Robin Larson, Sam Ringer, Scott Johnston, Shauna Kravec, Sheer El, Stanislav Fort, Tamera Lanham, Timothy Telleen-Lawton, Tom Conerly, Tom Henighan, Tristan Hume, Samuel Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown and Jared Kaplan · 2022
Earlier work this paper cites.
“Exploring Clusters of Research in Three Areas of AI Safety”, Center for Security and Emerging Technology, 2022
Helen Toner and Ashwin Acharya · 2022
Earlier work this paper cites.
“Taxonomy of Risks posed by Language Models”
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, Courtney Biles, Sasha Brown, Zac Kenton, Will Hawkins, Tom Stepleton, Abeba Birhane, Lisa Hendricks, Laura Rimell, William Isaac, Julia Haas, Sean Legassick, Geoffrey Irving and Iason Gabriel · 2022
Earlier work this paper cites.
“Discovering Language Model Behaviors with Model-Written Evaluations”, 2022
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Ben Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer and Jared Kaplan · 2022
Earlier work this paper cites.
“The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models”
Alexander Pan, Kush Bhatia and Jacob Steinhardt · 2022
Earlier work this paper cites.
“Goal Misgeneralization in Deep Reinforcement Learning” 162
Lauro Langosco, Jack Koch, Lee Sharkey, Jacob Pfau and David Krueger · 2022
Earlier work this paper cites.
“Goal Misgeneralization: Why Correct Specifications Aren’t Enough For Correct Goals”, 2022
Rohin Shah, Vikrant Varma, Ramana Kumar, Mary Phuong, Victoria Krakovna, Jonathan Uesato and Zac Kenton · 2022
Earlier work this paper cites.
“Adversarial Policies Beat Superhuman Go AIs”, 2022
Tony Wang, Adam Gleave, Tom Tseng, Kellin Pelrine, Nora Belrose, Joseph Miller, Michael Dennis, Yawen Duan, Viktor Pogrebniak, Sergey Levine and Stuart Russell · 2022
Cited alongside, same era.
“Emergent Abilities of Large Language Models”
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean and William Fedus · 2022
Cited alongside, same era.
“Chain-of-thought prompting elicits reasoning in large language models”
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le and Denny Zhou · 2022
Cited alongside, same era.
“Predictability and Surprise in Large Generative Models”
Deep Ganguli, Danny Hernandez, Liane Lovitt, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova Dassarma, Dawn Drain, Nelson Elhage, Sheer El, Stanislav Fort, Zac Hatfield-Dodds, Tom Henighan, Scott Johnston, Andy Jones, Nicholas Joseph, Jackson Kernian, Shauna Kravec, Ben Mann, Neel Nanda, Kamal Ndousse, Catherine Olsson, Daniela Amodei, Tom Brown, Jared Kaplan, Sam McCandlish, Christopher Olah, Dario Amodei and Jack Clark · 2022
Cited alongside, same era.
“Learning from human preferences” Accessed: 2023-9-15, https://openai.com/research/learning-from-human-preferences
Dario Amodei, Paul Christiano and Alex Ray · 2023
Closest in time.
“Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks”
Tilman Räuker, Anson Ho, Stephen Casper and Dylan Hadfield-Menell · 2023
Closest in time.
“AI capabilities can be significantly improved without expensive retraining”, 2023
Tom Davidson, Jean-Stanislas Denain, Pablo Villalobos and Guillem Bas · 2023
Closest in time.
“Model evaluation for extreme risks”, 2023
Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, Lewis Ho, Divya Siddarth, Shahar Avin, Will Hawkins, Been Kim, Iason Gabriel, Vijay Bolina, Jack Clark, Yoshua Bengio, Paul Christiano and Allan Dafoe · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Statement on AI Risk” Accessed: 2024-5-1, https://www.safe.ai/work/statement-on-ai-risk , 2023
2023
Cited alongside, same era.
“About” Accessed: 2023-9-15, https://www.deepmind.com/about
DeepMind · 2023
Cited alongside, same era.
“About” Accessed: 2023-9-15, https://openai.com/about
OpenAI · 2023
Cited alongside, same era.
“Trends in the Dollar Training Cost of Machine Learning Systems”, 2023
Ben Cottier · 2023
Cited alongside, same era.
“Trends in Machine Learning Hardware”, 2023
Marius Hobbhahn, Lennart Heim and Gökçe Aydos · 2023
Cited alongside, same era.
“Examples of AI Improving AI” Accessed: 2023-9-15, https://ai-improving-ai.safe.ai/
2023
Cited alongside, same era.
“GPT-4 Technical Report”, 2023
OpenAI · 2023
Cited alongside, same era.
“Harms from Increasingly Agentic Algorithmic Systems”
Alan Chan, Rebecca Salganik, Alva Markelius, Chris Pang, Nitarshan Rajkumar, Dmitrii Krasheninnikov, Lauro Langosco, Zhonghao He, Yawen Duan, Micah Carroll, Michelle Lin, Alex Mayhew, Katherine Collins, Maryam Molamohammadi, John Burden, Wanru Zhao, Shalaleh Rismani, Konstantinos Voudouris, Umang Bhatt, Adrian Weller, David Krueger and Tegan Maharaj · 2023
Cited alongside, same era.
Jérémy Scheurer, Mikita Balesni and Marius Hobbhahn · 2023
Closest in time.
Leonie Koessler and Jonas Schuett · 2023
Closest in time.
“The Bletchley Declaration by Countries Attending the AI Safety Summit, 1-2 November 2023”, 2023
AI Safety Summit · 2023
Closest in time.
“G7 Hiroshima Process on Generative Artificial Intelligence (AI)”
OECD · 2023
Closest in time.
“Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence”, 2023
The White House (US) · 2023
Closest in time.
“Interim Measures for Generative Artificial Intelligence Service Management” Accessed: 2024-2-12, http://www.cac.gov.cn/2023-07/13/c_1690898327029107.htm , 2023
Cyberspace Administration of China · 2023
Closest in time.
“A pro-innovation approach to AI regulation” Accessed: 2024-2-12
Department of State for Science, Innovation and Technology (UK) · 2023
Closest in time.
“International Institutions for Advanced AI”, 2023
Lewis Ho, Joslyn Barnhart, Robert Trager, Yoshua Bengio, Miles Brundage, Allison Carnegie, Rumman Chowdhury, Allan Dafoe, Gillian Hadfield, Margaret Levi and Duncan Snidal · 2023
Closest in time.
“International Governance of Civilian AI: A Jurisdictional Certification Approach”, https://cdn.governance.ai/International_Governance_of_Civilian_AI_OMS.pdf , 2023
Robert Trager, Ben Harack, Anka Reuel, Allison Carnegie, Lennart Heim, Lewis Ho, Sarah Kreps, Ranjit Lall, Owen Larter, Seán ÓÉigeartaigh, Simon Staffell and José Villalobos · 2023
Closest in time.
“Frontier AI Regulation: Managing Emerging Risks to Public Safety”, 2023
Markus Anderljung, Joslyn Barnhart, Anton Korinek, Jade Leung, Cullen O’Keefe, Jess Whittlestone, Shahar Avin, Miles Brundage, Justin Bullock, Duncan Cass-Beggs, Ben Chang, Tantum Collins, Tim Fist, Gillian Hadfield, Alan Hayes, Lewis Ho, Sara Hooker, Eric Horvitz, Noam Kolt, Jonas Schuett, Yonadav Shavit, Divya Siddarth, Robert Trager and Kevin Wolf · 2023
Closest in time.
“Auditing large language models: a three-layered approach”
Jakob Mökander, Jonas Schuett, Hannah Kirk and Luciano Floridi · 2023
Closest in time.
“SMP12. Safety Case and Safety Case Report” Accessed: 2024-2-12, https://www.asems.mod.uk/guidance/posms/smp12 , 2023
2023
Closest in time.
“ISO/IEC 23894:2023 Standard on Information technology — Artificial intelligence — Guidance on risk management”, 2023
Iso/iec · 2023
Closest in time.
“General Purpose AI Poses Serious Risks, Should Not Be Excluded From the EU’s AI Act — Policy Brief” Accessed: 2023-9-15, https://ainowinstitute.org/publication/gpai-is-high-risk-should-not-be-excluded-from-eu-ai-act
AI Now Institute · 2023
Closest in time.
“Towards best practices in AGI safety and governance: A survey of expert opinion”, 2023
Jonas Schuett, Noemi Dreksler, Markus Anderljung, David McCaffary, Lennart Heim, Emma Bluemke and Ben Garfinkel · 2023
Closest in time.
“Regulatory Markets: The Future of AI Governance”, 2023
Gillian Hadfield and Jack Clark · 2023
Closest in time.
“AI safety – ETO Research Almanac” Accessed: 2024-2-12, https://almanac.eto.tech/topics/ai-safety/
Emerging Technology Observatory · 2024
Closest in time.
“The alignment problem from a deep learning perspective”
Richard Ngo, Lawrence Chan and Sören Mindermann · 2024
Closest in time.
“Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training”, 2024
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel Ziegler, Tim Maxwell, Newton Cheng, Adam Jermyn, Amanda Askell, Ansh Radhakrishnan, Cem Anil, David Duvenaud, Deep Ganguli, Fazl Barez, Jack Clark, Kamal Ndousse, Kshitij Sachan, Michael Sellitto, Mrinank Sharma, Nova DasSarma, Roger Grosse, Shauna Kravec, Yuntao Bai, Zachary Witten, Marina Favaro, Jan Brauner, Holden Karnofsky, Paul Christiano, Samuel Bowman, Logan Graham, Jared Kaplan, Sören Mindermann, Ryan Greenblatt, Buck Shlegeris, Nicholas Schiefer and Ethan Perez · 2024
Closest in time.
“Regulating advanced artificial agents”
Michael Cohen, Noam Kolt, Yoshua Bengio, Gillian Hadfield and Stuart Russell · 2024
Closest in time.
“Self-Discover: Large Language Models Self-Compose Reasoning Structures”, 2024
Pei Zhou, Jay Pujara, Xiang Ren, Xinyun Chen, Heng-Tze Cheng, Quoc Le, Ed Chi, Denny Zhou, Swaroop Mishra and Huaixiu Zheng · 2024
Closest in time.
“The Operational Risks of AI in Large-Scale Biological Attacks: Results of a Red-Team Study”
Christopher Mouton, Caleb Lucas and Ella Guest · 2024
Closest in time.
“EU AI Act” Accessed: 2024-NA-NA, https://artificialintelligenceact.eu/the-act/ , 2024
European Union · 2024
Closest in time.
“Responsible Reporting for Frontier AI Development”, 2024
Noam Kolt, Markus Anderljung, Joslyn Barnhart, Asher Brass, Kevin Esvelt, Gillian Hadfield, Lennart Heim, Mikel Rodriguez, Jonas Sandbrink and Thomas Woodside · 2024
Closest in time.
“Black-Box Access is Insufficient for Rigorous AI Audits”, 2024
Stephen Casper, Carson Ezell, Charlotte Siegmann, Noam Kolt, Taylor Curtis, Benjamin Bucknall, Andreas Haupt, Kevin Wei, Jérémy Scheurer, Marius Hobbhahn, Lee Sharkey, Satyapriya Krishna, Marvin Von, Silas Alberti, Alan Chan, Qinyi Sun, Michael Gerovitch, David Bau, Max Tegmark, David Krueger and Dylan Hadfield-Menell · 2024
Closest in time.
“Evaluating Frontier Models for Dangerous Capabilities”, 2024
Mary Phuong, Matthew Aitchison, Elliot Catt, Sarah Cogan, Alexandre Kaskasoli, Victoria Krakovna, David Lindner, Matthew Rahtz, Yannis Assael, Sarah Hodkinson, Heidi Howard, Tom Lieberum, Ramana Kumar, Maria Raad, Albert Webson, Lewis Ho, Sharon Lin, Sebastian Farquhar, Marcus Hutter, Gregoire Deletang, Anian Ruoss, Seliem El-Sayed, Sasha Brown, Anca Dragan, Rohin Shah, Allan Dafoe and Toby Shevlane · 2024
Closest in time.
“Safety Cases: How to Justify the Safety of Advanced AI Systems”, 2024
Joshua Clymer, Nick Gabrieli, David Krueger and Thomas Larsen · 2024
Closest in time.