Fetching the paper…
Reading the bibliography…
Proposals for safe AGI systems are typically made at the level of frameworks, specifying how the components of the proposed system should be trained and interact with each other.
Influence diagrams
Ronald A Howard and James E Matheson · 1984
Earlier work this paper cites.
The Intentional Stance
Daniel Dennett · 1987
Earlier work this paper cites.
Probabilistic evaluation of counterfactual queries
Alexander Balke and Judea Pearl · 1994
Earlier work this paper cites.
Multi-agent influence diagrams for representing and solving games
Daphne Koller and Brian Milch · 2003
Earlier work this paper cites.
Godel machines: Self-referential universal problem solvers making provably optimal self-improvements
Jürgen Schmidhuber · 2007
Earlier work this paper cites.
Causality: Models, Reasoning, and Inference
Judea Pearl · 2009
Earlier work this paper cites.
Self-modification and mortality in artificial agents
Laurent Orseau and Mark Ring · 2011
Earlier work this paper cites.
Model-based utility functions
Bill Hibbard · 2012
Earlier work this paper cites.
Quantifying causal emergence shows that macro can beat micro
Erik Hoel, Larissa Albantakis, and Giulio Tononi · 2013
Earlier work this paper cites.
Superintelligence: Paths, Dangers, Strategies
Nick Bostrom · 2014
Cited alongside, same era.
Act-based agents, 2015
Paul Christiano · 2015
Cited alongside, same era.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Cited alongside, same era.
Why tool AIs want to be agent AIs, 2016
Gwern Branwen · 2016
Cited alongside, same era.
Self-modification of policy and utility function in rational agents
Tom Everitt, Daniel Filan, Mayank Daswani, and Marcus Hutter · 2016
Cited alongside, same era.
Cooperative inverse reinforcement learning
Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell · 2016
Cited alongside, same era.
AGI safety literature review
Tom Everitt, Gary Lea, and Marcus Hutter · 2018
Later among the works it cites.
Towards Safe Artificial General Intelligence
Tom Everitt · 2018
Later among the works it cites.
Geoffrey Irving, Paul Christiano, and Dario Amodei · 2018
Later among the works it cites.
Scalable agent alignment via reward modeling: a research direction
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg · 2018
Later among the works it cites.
Agents and devices: A relative definition of agency
Laurent Orseau, Simon McGregor McGill, and Shane Legg · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Good and safe uses of AI oracles
Stuart Armstrong · 2017
Cited alongside, same era.
Learning to reinforcement learn
Jane Wang, Zeb Kurth-Nelson, Hubert Soyer, Joel Leibo, Dhruva Tirumala, Rémi Munos, Charles Blundell, Dharshan Kumaran, and Matt Botvinick · 2017
Cited alongside, same era.
Supervising strong learners by amplifying weak experts
Paul Christiano, Buck Shlegeris, and Dario Amodei · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Reframing superintelligence: Comprehensive AI services as general intelligence
K Eric Drexler · 2019
Closest in time.
Understanding agent incentives using causal influence diagrams. Part I: Single action settings
Tom Everitt, Pedro Ortega, Elizabeth Barnes, and Shane Legg · 2019
Closest in time.