Fetching the paper…
Reading the bibliography…
Some researchers speculate that intelligent reinforcement learning (RL) agents would be incentivized to seek resources and power in pursuit of their objectives.
On the set of optimal policies in discrete dynamic programming
Steven A Lippman · 1968
Earlier work this paper cites.
Turnpike theory
Lionel W McKenzie · 1976
Earlier work this paper cites.
Composing functions to speed up reinforcement learning in a changing world
Chris Drummond · 1998
Earlier work this paper cites.
Reinforcement learning: an introduction
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y. Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Structure in the space of value functions
David Foster and Peter Dayan · 2002
Earlier work this paper cites.
Q-cut—dynamic discovery of sub-goals in reinforcement learning
Ishai Menache, Shie Mannor, and Nahum Shimkin · 2002
Earlier work this paper cites.
Dual representations for dynamic programming and reinforcement learning
Tao Wang, Michael Bowling, and Dale Schuurmans · 2007
Earlier work this paper cites.
The basic AI drives, 2008
Stephen Omohundro · 2008
Earlier work this paper cites.
Stable dual dynamic programming
Tao Wang, Michael Bowling, Dale Schuurmans, and Daniel J Lizotte · 2008
Earlier work this paper cites.
Artificial intelligence: a modern approach
Stuart J Russell and Peter Norvig · 2009
Earlier work this paper cites.
Robust policy computation in reward-uncertain MDPs using nondominated policies
Kevin Regan and Craig Boutilier · 2010
Earlier work this paper cites.
Campbell Biology
J.B. Reece and N.A. Campbell · 2011
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Richard S Sutton, Joseph Modayil, Michael Delp, Thomas Degris, Patrick M Pilarski, Adam White, and Doina Precup · 2011
Cited alongside, same era.
The superintelligent will: Motivation and instrumental rationality in advanced artificial agents
Nick Bostrom · 2012
Cited alongside, same era.
Superintelligence
Nick Bostrom · 2014
Cited alongside, same era.
Markov decision processes: Discrete stochastic dynamic programming
Martin L Puterman · 2014
Cited alongside, same era.
Empowerment–an introduction
Christoph Salge, Cornelius Glackin, and Daniel Polani · 2014
Cited alongside, same era.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Cited alongside, same era.
Incorrigibility in the CIRL framework
Ryan Carey · 2018
Later among the works it cites.
New and surprising ways to be mean
Christian Guckelsberger, Christoph Salge, and Julian Togelius · 2018
Later among the works it cites.
Large-scale study of curiosity-driven learning
Yuri Burda, Harri Edwards, Deepak Pathak, Amos Storkey, Trevor Darrell, and Alexei A. Efros · 2019
Closest in time.
Don’t fear the Terminator, September 2019
Yann LeCun and Anthony Zador · 2019
Closest in time.
Human compatible: Artificial intelligence and the problem of control
Stuart Russell · 2019
Closest in time.
Power and technology: a philosophical and ethical analysis
Faridun Sattarov · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Corrigibility
Nate Soares, Benja Fallenstein, Stuart Armstrong, and Eliezer Yudkowsky · 2015
Cited alongside, same era.
Formalizing convergent instrumental goals
Tsvi Benson-Tilsen and Nate Soares · 2016
Cited alongside, same era.
Intrinsically motivated general companion NPCs via coupled empowerment maximisation
Christian Guckelsberger, Christoph Salge, and Simon Colton · 2016
Cited alongside, same era.
The off-switch game
Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell · 2017
Cited alongside, same era.
Should robots be obedient?
Smitha Milli, Dylan Hadfield-Menell, Anca Dragan, and Stuart Russell · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell · 2017
Cited alongside, same era.
Various · 2019
Closest in time.
AvE: Assistance via empowerment
Yuqing Du, Stas Tiomkin, Emre Kiciman, Daniel Polani, Pieter Abbeel, and Anca Dragan · 2020
Closest in time.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Closest in time.
The foundations, benefits, and possible existential threat of AI, June 2020
Steven Pinker and Stuart Russell · 2020
Closest in time.
Conservative agency via attainable utility preservation
Alexander Matt Turner, Dylan Hadfield-Menell, and Prasad Tadepalli · 2020
Closest in time.
Why AI is harder than we think
Melanie Mitchell · 2021
Closest in time.
Reward is not the optimization target, 2022
Alexander Matt Turner · 2022
Closest in time.