Fetching the paper…
Reading the bibliography…
While it is still unclear if agents with Artificial General Intelligence (AGI) could ever be built, we can already use mathematical models to investigate potential safety systems for these agents.
Craig Boutilier, Thomas Dean, and Steve Hanks, Decision-theoretic planning: Structural assumptions and computational leverage , J. Artif. Int. Res. 11
1999
Earlier work this paper cites.
Stephen M Omohundro, The basic AI drives , AGI, vol. 171, 2008, pp. 483–492
2008
Earlier work this paper cites.
Judea Pearl, Causality , Cambridge university press, 2009
2009
Earlier work this paper cites.
Ross Shachter and David Heckerman, Pearl causality and the value of control , Heuristics, Probability, and Causality: A Tribute to Judea Pearl, College Publications, London, 2010, pp. 431–447
2010
Earlier work this paper cites.
Stuart Armstrong, Motivated value selection for artificial agents , Workshops at the Twenty-Ninth AAAI Conference on Artificial Intelligence, 2015
2015
Earlier work this paper cites.
Javier Garcıa and Fernando Fernández, A comprehensive survey on safe reinforcement learning , Journal of Machine Learning Research 16
2015
Earlier work this paper cites.
Nate Soares, Benja Fallenstein, Stuart Armstrong, and Eliezer Yudkowsky, Corrigibility , Workshops at the Twenty-Ninth AAAI Conference on Artificial Intelligence, 2015
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
Laurent Orseau and Stuart Armstrong, Safely interruptible agents , Proceedings of the Thirty-Second Conference on Uncertainty in Artificial Intelligence, AUAI Press, 2016, pp. 557–566
2016
Cited alongside, same era.
2017
Cited alongside, same era.
Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell, The off-switch game , Workshops at the Thirty-First AAAI Conference on Artificial Intelligence, 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2019
Later among the works it cites.
2019
Later among the works it cites.
Koen Holtman, Corrigibility with utility preservation , arXiv:1908.01695 (2019)
2019
Later among the works it cites.
Stuart Jonathan Russell, Human compatible: Artificial intelligence and the problem of control , Penguin Random House, 2019
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tom Everitt, Gary Lea, and Marcus Hutter, AGI safety literature review , Proceedings of the 27th International Joint Conference on Artificial Intelligence, AAAI Press, 2018, pp. 5441–5449
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Abram Demski and Scott Garrabrant, Embedded agency , arXiv:1902.09469 (2019)
2019
Cited alongside, same era.
2020
Closest in time.
Koen Holtman, Towards AGI agent safety by iteratively improving the utility function , Proceedings of the 13th International Conference on Artificial General Intelligence (AGI-20). Lecture Notes in Computer Science, vol 12177, Springer, 2020
2020
Closest in time.
Alexander Matt Turner, Dylan Hadfield-Menell, and Prasad Tadepalli, Conservative agency via attainable utility preservation , Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 2020, pp. 385–391
2020
Closest in time.