Fetching the paper…
Reading the bibliography…
`Indifference' refers to a class of methods used to control reward based agents.
Reinforcement Learning: An Introduction
R. Sutton and A.G. Barto · 1998
Earlier work this paper cites.
The basic ai drives
Stephen M Omohundro · 2008
Earlier work this paper cites.
Utility indifference
Stuart Armstrong · 2010
Earlier work this paper cites.
Marcus Hutter · 2012
Earlier work this paper cites.
Superintelligence: Paths, dangers, strategies
Nick Bostrom · 2014
Earlier work this paper cites.
Utility indifference and infinite improbability drives
Benja Fallenstein · 2014
Earlier work this paper cites.
Motivated value selection for artificial agents
Stuart Armstrong · 2015
Earlier work this paper cites.
Learning the preferences of ignorant, inconsistent agents
Owain Evans, Andreas Stuhlmüller, and Noah D. Goodman · 2015
Earlier work this paper cites.
Corrigibility
Nate Soares, Benja Fallenstein, Eliezer Yudkowsky, and Stuart Armstrong · 2015
Earlier work this paper cites.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul F. Christiano, John Schulman, and Dan Mané · 2016
Cited alongside, same era.
Self-modification of policy and utility function in rational agents
Tom Everitt, Daniel Filan, Mayank Daswani, and Marcus Hutter · 2016
Cited alongside, same era.
Dylan Hadfield-Menell, Anca D. Dragan, Pieter Abbeel, and Stuart J. Russell · 2016
Cited alongside, same era.
Safely interruptible agents
Laurent Orseau and Stuart Armstrong · 2016
Cited alongside, same era.
Research priorities for robust and beneficial artificial intelligence
Stuart J. Russell, Daniel Dewey, and Max Tegmark · 2016
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Closest in time.
When will AI exceed human performance? evidence from AI experts
Katja Grace, John Salvatier, Allan Dafoe, Baobao Zhang, and Owain Evans · 2017
Closest in time.
Inverse reward design
Dylan Hadfield-Menell, Smitha Milli, Stuart J Russell, Pieter Abbeel, and Anca Dragan · 2017
Closest in time.
Jan Leike, Miljan Martic, Victoria Krakovna, Pedro A. Ortega, Tom Everitt, Andrew Lefrancq, Laurent Orseau, and Shane Legg · 2017
Closest in time.
Smitha Milli, Dylan Hadfield-Menell, Anca D. Dragan, and Stuart J. Russell · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Graying the black box: Understanding dqns
Tom Zahavy, Nir Ben-Zrihem, and Shie Mannor · 2016
Cited alongside, same era.
Good and safe uses of AI oracles
Stuart Armstrong · 2017
Cited alongside, same era.
On the promotion of safe and socially beneficial artificial intelligence
Seth D. Baum · 2017
Cited alongside, same era.
Tom Everitt, Gary Lea, and Marcus Hutter
Cited in the paper.
Enter the matrix: A virtual world approach to safely interruptable autonomous systems
Mark O. Riedl and Brent Harrison · 2017
Closest in time.
Counterfactual equivalence for pomdps, and underlying deterministic environments
Stuart Armstrong · 2018
Closest in time.
Agents manipulating their own learning process
Stuart Armstrong, Jan Leike, Laurent Orseau, and Shane Legg · 2018
Closest in time.