Fetching the paper…
Reading the bibliography…
It is clear that one of the primary tools we can use to mitigate the potential risk from a misbehaving AI system is the ability to turn the system off.
Can digital machines think?
Alan M. Turing · 1951
Earlier work this paper cites.
On the folly of rewarding a, while hoping for b
Steven Kerr · 1975
Earlier work this paper cites.
Incentives in organizations
Robert Gibbons · 1998
Earlier work this paper cites.
The basic AI drives
Stephen M. Omohundro · 2008
Earlier work this paper cites.
Cognition and incomplete contracts
Jean Tirole · 2009
Earlier work this paper cites.
Artificial Intelligence: A Modern Approach
Stuart Russell and Peter Norvig · 2010
Cited alongside, same era.
A decision-theoretic model of assistance
Alan Fern, Sriraam Natarajan, Kshitij Judah, and Prasad Tadepalli · 2014
Cited alongside, same era.
Here’s what Facebook’s artificial intelligence expert thinks about the future
Guia Marie Del Prado · 2015
Cited alongside, same era.
Are super intelligent computers really a threat to humanity?
ITIF · 2015
Cited alongside, same era.
Corrigibility
Nate Soares, Benja Fallenstein, Stuart Armstrong, and Eliezer Yudkowsky · 2015
Later among the works it cites.
Cooperative inverse reinforcement learning
Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell · 2016
Closest in time.
Safely interruptible agents
Laurent Orseau and Stuart Armstrong · 2016
Closest in time.
Should we fear supersmart robots?
Stuart Russell · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…