Fetching the paper…
Reading the bibliography…
How can humans stay in control of advanced artificial intelligence systems? One proposal is corrigibility, which requires the agent to follow the instructions of a human overseer, without inappropriately influencing them.
Influence diagrams for causal modelling and inference
A Philip Dawid · 2002
Earlier work this paper cites.
The basic AI drives
Stephen M Omohundro · 2008
Earlier work this paper cites.
Causality
Judea Pearl · 2009
Earlier work this paper cites.
Utility indifference
Stuart Armstrong · 2010
Earlier work this paper cites.
Superintelligence: Paths, dangers, strategies., 2014a
Nick Bostrom · 2014
Earlier work this paper cites.
The multi-slot framework: A formal model for multiple, copiable AIs
Laurent Orseau · 2014
Earlier work this paper cites.
Research priorities for robust and beneficial artificial intelligence
Stuart Russell, Daniel Dewey, and Max Tegmark · 2015
Earlier work this paper cites.
Corrigibility
Nate Soares, Benja Fallenstein, Stuart Armstrong, and Eliezer Yudkowsky · 2015
Earlier work this paper cites.
Cooperative inverse reinforcement learning
Dylan Hadfield-Menell, Stuart J Russell, Pieter Abbeel, and Anca Dragan · 2016
Earlier work this paper cites.
Death and suicide in universal artificial intelligence
Jarryd Martin, Tom Everitt, and Marcus Hutter · 2016
Earlier work this paper cites.
Safely interruptible agents
Laurent Orseau and Stuart Armstrong · 2016
Earlier work this paper cites.
Two problems with causal-counterfactual utility indifference
Jessica Taylor · 2016
Earlier work this paper cites.
’indifference’methods for managing agent rewards
Stuart Armstrong and Xavier O’Rourke · 2017
Earlier work this paper cites.
Corrigibility
Paul Christiano · 2017
Cited alongside, same era.
The off-switch game
Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell · 2017
Cited alongside, same era.
Jan Leike, Miljan Martic, Victoria Krakovna, Pedro A Ortega, Tom Everitt, Andrew Lefrancq, Laurent Orseau, and Shane Legg · 2017
Cited alongside, same era.
Should robots be obedient?
Smitha Milli, Dylan Hadfield-Menell, Anca Dragan, and Stuart Russell · 2017
Cited alongside, same era.
Agent foundations for aligning machine intelligence with human interests: a technical research agenda
Nate Soares and Benya Fallenstein · 2017
Cited alongside, same era.
Incorrigibility in the CIRL framework
Ryan Carey · 2018
Cited alongside, same era.
Conservative agency via attainable utility preservation
Alexander M. Turner, Dylan Hadfield-Menell, and Prasad Tadepalli · 2020
Later among the works it cites.
Darpa’s explainable AI (XAI) program: A retrospective
David Gunning, Eric Vorm, Jennifer Yunyan Wang, and Matt Turek · 2021
Later among the works it cites.
Human-compatible artificial intelligence
Stuart Russell · 2021
Later among the works it cites.
Optimal policies tend to seek power
Alexander Matt Turner, Logan Smith, Rohin Shah, Andrew Critch, and Prasad Tadepalli · 2021
Later among the works it cites.
Definitions of intent suitable for algorithms
Hal Ashton · 2022
Later among the works it cites.
Constitutional AI: Harmlessness from AI feedback, 2022
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Kamile Lukosuite, Liane Lovitt, Michael Sellitto, Nelson Elhage, Nicholas Schiefer, Noemi Mercado, Nova DasSarma, Robert Lasenby, Robin Larson, Sam Ringer, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Tamera Lanham, Timothy Telleen-Lawton, Tom Conerly, Tom Henighan, Tristan Hume, Samuel R. Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards formal definitions of blameworthiness, intention, and moral responsibility
Joseph Y Halpern and Max Kleiman-Weiner · 2018
Cited alongside, same era.
AI safety via debate, 2018
Geoffrey Irving, Paul Christiano, and Dario Amodei · 2018
Cited alongside, same era.
Scalable agent alignment via reward modeling: a research direction, 2018
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg · 2018
Cited alongside, same era.
Corrigibility with utility preservation, 2020
Koen Holtman · 2020
Cited alongside, same era.
Avoiding side effects by considering future tasks
Victoria Krakovna, Laurent Orseau, Richard Ngo, Miljan Martic, and Shane Legg · 2020
Cited alongside, same era.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Cited alongside, same era.
Later among the works it cites.
Path-specific objectives for safer agent incentives
Sebastian Farquhar, Ryan Carey, and Tom Everitt · 2022
Later among the works it cites.
Human-centred mechanism design with democratic AI
Raphael Koster, Jan Balaguer, Andrea Tacchetti, Ari Weinstein, Tina Zhu, Oliver Hauser, Duncan Williams, Lucy Campbell-Gillingham, Phoebe Thacker, Matthew Botvinick, and Christopher Summerfield · 2022
Later among the works it cites.
Jonathan G Richens, Rory Beard, and Daniel H Thompson · 2022
Later among the works it cites.
Equivalence and synthesis of causal models
Thomas S Verma and Judea Pearl · 2022
Later among the works it cites.
Problem of fully updated deference
Arbital · 2023
Closest in time.
What Failure Looks Like
Paul Christiano · 2023
Closest in time.