Fetching the paper…
Reading the bibliography…
If AI systems match or exceed human capabilities on a wide range of tasks, it may become difficult for humans to efficiently judge their actions -- making it hard to use human feedback to steer them towards desirable traits.
Risks from learned optimization in advanced machine learning systems, 2021
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant · 1906
Earlier work this paper cites.
Defence standard 00-56 issue 4: Safety management requirements for defence systems
UK Ministry of Defence · 2007
Earlier work this paper cites.
Building blocks for assurance cases
Robin Bloomfield and Kateryna Netkachova · 2014
Earlier work this paper cites.
Superintelligence: Paths, Dangers, Strategies
Nick Bostrom · 2014
Earlier work this paper cites.
Concrete problems in AI safety, 2016
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts, 2018
Paul Christiano, Buck Shlegeris, and Dario Amodei · 2018
Earlier work this paper cites.
Geoffrey Irving, Paul Christiano, and Dario Amodei · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction, 2018
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg · 2018
Earlier work this paper cites.
Human Compatible: Artificial Intelligence and the Problem of Control
Stuart Russell · 2019
Earlier work this paper cites.
The Alignment Problem: Machine Learning and Human Values
Brian Christian · 2020
Earlier work this paper cites.
Measuring progress on scalable oversight for large language models, 2022
Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamilė Lukošiūtė, Amanda Askell, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Christopher Olah, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Jackson Kernion, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Liane Lovitt, Nelson Elhage, Nicholas Schiefer, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Robin Larson, Sam McCandlish, Sandipan Kundu, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Ben Mann, and Jared Kaplan · 2022
Earlier work this paper cites.
Goal misgeneralization in deep reinforcement learning
Lauro Langosco Di Langosco, Jack Koch, Lee D Sharkey, Jacob Pfau, and David Krueger · 2022
Earlier work this paper cites.
The effects of reward misspecification: Mapping and mitigating misaligned models
Alexander Pan, Kush Bhatia, and Jacob Steinhardt · 2022
Earlier work this paper cites.
QuALITY: Question answering with long input texts, yes!
Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He, and Samuel Bowman · 2022
Earlier work this paper cites.
Goal misgeneralization: Why correct specifications aren’t enough for correct goals, 2022
Rohin Shah, Vikrant Varma, Ramana Kumar, Mary Phuong, Victoria Krakovna, Jonathan Uesato, and Zac Kenton · 2022
Earlier work this paper cites.
Scalable AI safety via doubly-efficient debate, 2023
Jonah Brown-Cohen, Geoffrey Irving, and Georgios Piliouras · 2023
Earlier work this paper cites.
Power-seeking can be probable and predictive for trained agents, 2023
Victoria Krakovna and Janos Kramar · 2023
Earlier work this paper cites.
Debate helps supervise unreliable experts, 2023
Julian Michael, Salsabila Mahdi, David Rein, Jackson Petty, Julien Dirani, Vishakh Padmakumar, and Samuel R. Bowman · 2023
Cited alongside, same era.
Responsible scaling policy
Anthropic · 2024
Cited alongside, same era.
Training language models to win debates with self-play improves judge accuracy, 2024
Samuel Arnesen, David Rein, and Julian Michael · 2024
Cited alongside, same era.
Towards evaluations-based safety cases for AI scheming, 2024
Mikita Balesni, Marius Hobbhahn, David Lindner, Alexander Meinke, Tomek Korbak, Joshua Clymer, Buck Shlegeris, Jérémy Scheurer, Charlotte Stix, Rusheb Shah, Nicholas Goldowsky-Dill, Dan Braun, Bilal Chughtai, Owain Evans, Daniel Kokotajlo, and Lucius Bushnaq · 2024
Cited alongside, same era.
Managing extreme AI risks amid rapid progress
Yoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Trevor Darrell, Yuval Noah Harari, Ya-Qin Zhang, Lan Xue, Shai Shalev-Shwartz, Gillian Hadfield, Jeff Clune, Tegan Maharaj, Frank Hutter, Atılım Güneş Baydin, Sheila McIlraith, Qiqi Gao, Ashwin Acharya, David Krueger, Anca Dragan, Philip Torr, Stuart Russell, Daniel Kahneman, Jan Brauner, and Sören Mindermann · 2024
Studying small language models with susceptibilities, 2025
Garrett Baker, George Wang, Jesse Hoogland, and Daniel Murfet · 2025
Closest in time.
Debate update: Obfuscated arguments problem
Beth Barnes · 2025
Closest in time.
Writeup: Progress on AI safety via debate
Beth Barnes and Paul Christiano · 2025
Closest in time.
International AI safety report
Yoshua Bengio, Sören Mindermann, Daniel Privitera, Tamay Besiroglu, Rishi Bommasani, Stephen Casper, Yejin Choi, Philip Fox, Ben Garfinkel, Danielle Goldfarb, Hoda Heidari, Anson Ho, Sayash Kapoor, Leila Khalatbari, Shayne Longpre, Sam Manning, Vasilios Mavroudis, Mantas Mazeika, Julian Michael, Jessica Newman, Kwan Yee Ng, Chinasa T. Okolo, Deborah Raji, Girish Sastry, Elizabeth Seger, Theodora Skeadas, Tobin South, Emma Strubell, Florian Tramèr, Lucia Velasco, Nicole Wheeler, Daron Acemoglu, Olubayo Adekanmbi, David Dalrymple, Thomas G. Dietterich, Edward W. Felten, Pascale Fung, Pierre-Olivier Gourinchas, Fredrik Heintz, Geoffrey Hinton, Nick Jennings, Andreas Krause, Susan Leavy, Percy Liang, Teresa Ludermir, Vidushi Marda, Helen Margetts, John McDermid, Jane Munga, Arvind Narayanan, Alondra Nelson, Clara Neppel, Alice Oh, Gopal Ramchurn, Stuart Russell, Marietje Schaake, Bernhard Schölkopf, Dawn Song, Alvaro Soto, Lee Tiedrich, Gaël Varoquaux, Andrew Yao, Ya-Qin Zhang, Olubunmi Ajala, Fahad Albalawi, Marwan Alserkal, Guillaume Avrin, Christian Busch, André Carlos Ponce de Leon Ferreira de Carvalho, Bronwyn Fox, Amandeep Singh Gill, Ahmet Halit Hatip, Juha Heikkilä, Chris Johnson, Gill Jolly, Ziv Katzir, Saif M. Khan, Hiroaki Kitano, Antonio Krüger, Kyoung Mu Lee, Dominic Vincent Ligot, José Ramón López Portillo, Oleksii Molchanovskyi, Andrea Monti, Nusu Mwamanzi, Mona Nemer, Nuria Oliver, Raquel Pezoa Rivera, Balaraman Ravindran, Hammam Riza, Crystal Rugege, Ciarán Seoighe, Jerry Sheehan, Haroon Sheikh, Denise Wong, and Yi Zeng · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Safety cases for frontier AI, 2024
Marie Davidsen Buhl, Gaurav Sett, Leonie Koessler, Jonas Schuett, and Markus Anderljung · 2024
Cited alongside, same era.
Safety cases: How to justify the safety of advanced AI systems, 2024
Joshua Clymer, Nick Gabrieli, David Krueger, and Thomas Larsen · 2024
Cited alongside, same era.
Safety case template for frontier AI: A cyber inability argument, 2024
Arthur Goemans, Marie Davidsen Buhl, Jonas Schuett, Tomek Korbak, Jessica Wang, Benjamin Hilton, and Geoffrey Irving · 2024
Cited alongside, same era.
Alignment faking in large language models, 2024
Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid, Sam Marks, Johannes Treutlein, Tim Belonax, Jack Chen, David Duvenaud, Akbir Khan, Julian Michael, Sören Mindermann, Ethan Perez, Linda Petrini, Jonathan Uesato, Jared Kaplan, Buck Shlegeris, Samuel R. Bowman, and Evan Hubinger · 2024
Cited alongside, same era.
Safety cases at AISI
Geoffrey Irving · 2024
Cited alongside, same era.
On scalable oversight with weak llms judging strong llms, 2024
Zachary Kenton, Noah Y. Siegel, János Kramár, Jonah Brown-Cohen, Samuel Albanie, Jannis Bulian, Rishabh Agarwal, David Lindner, Yunhao Tang, Noah D. Goodman, and Rohin Shah · 2024
Cited alongside, same era.
Debating with more persuasive llms leads to more truthful answers, 2024
Akbir Khan, John Hughes, Dan Valentine, Laura Ruis, Kshitij Sachan, Ansh Radhakrishnan, Edward Grefenstette, Samuel R. Bowman, Tim Rocktäschel, and Ethan Perez · 2024
Cited alongside, same era.
Closest in time.
Low-Stakes alignment
Paul Christiano · 2025
Closest in time.
Frontier Safety Framework
Google DeepMind · 2025
Closest in time.
Notes on countermeasures for exploration hacking (aka sandbagging)
Ryan Greenblatt · 2025
Closest in time.
Safety cases: A scalable approach to frontier AI safety, 2025
Benjamin Hilton, Marie Davidsen Buhl, Tomek Korbak, and Geoffrey Irving · 2025
Closest in time.
Formal verification, heuristic explanations and surprise accounting, June 2024
Jacob Hilton · 2025
Closest in time.
Eliciting bad contexts
Geoffrey Irving, Joseph Isaac Bloom, and Tomek Korbak · 2025
Closest in time.
AI alignment: A comprehensive survey, 2025
Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Lukas Vierling, Donghai Hong, Jiayi Zhou, Zhaowei Zhang, Fanzhi Zeng, Juntao Dai, Xuehai Pan, Kwan Yee Ng, Aidan O’Gara, Hua Xu, Brian Tse, Jie Fu, Stephen McAleer, Yaodong Yang, Yizhou Wang, Song-Chun Zhu, Yike Guo, and Wen Gao · 2025
Closest in time.
A sketch of an AI control safety case, 2025
Tomek Korbak, Joshua Clymer, Benjamin Hilton, Buck Shlegeris, and Geoffrey Irving · 2025
Closest in time.
The alignment problem from a deep learning perspective, 2025
Richard Ngo, Lawrence Chan, and Sören Mindermann · 2025
Closest in time.
Preparedness framework
OpenAI · 2025
Closest in time.
An approach to technical AGI safety and security
Rohin Shah, Alex Irpan, Alexander Matt Turner, Anna Wang, Arthur Conmy, David Lindner, Jonah Brown-Cohen, Lewis Ho, Neel Nanda, Raluca Ada Popa, Rishub Jain, Rory Greig, Samuel Albanie, Scott Emmons, Sebastian Farquhar, Sébastien Krier, Senthooran Rajamanoharan, Sophie Bridgers, Tobi Ijitoye, Tom Everitt, Victoria Krakovna, Vikrant Varma, Vladimir Mikulik, Zachary Kenton, Dave Orr, Shane Legg, Noah Goodman, Allan Dafoe, Four Flynn, and Anca Dragan · 2025
Closest in time.
Defining and characterizing reward hacking, 2025
Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, and David Krueger · 2025
Closest in time.
Frontier AI safety commitments, AI seoul summit 2024
UK and Republic of Korea · 2025
Closest in time.