Fetching the paper…
Reading the bibliography…
If capable AI agents are generally incentivized to seek power in service of the objectives we specify for them, then these systems will pose enormous risks, in addition to enormous benefits.
Risks from learned optimization in advanced machine learning systems, 2019
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant · 1906
Earlier work this paper cites.
Rational choice and the structure of the environment
Herbert A Simon · 1956
Earlier work this paper cites.
Reinforcement learning: an introduction
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Goal inference as inverse planning
Chris L Baker, Joshua B Tenenbaum, and Rebecca R Saxe · 2007
Earlier work this paper cites.
Superintelligence
Nick Bostrom · 2014
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Martin L Puterman · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Massively parallel methods for deep reinforcement learning
Arun Nair, Praveen Srinivasan, Sam Blackwell, Cagdas Alcicek, Rory Fearon, Alessandro De Maria, Vedavyas Panneershelvam, Mustafa Suleyman, Charles Beattie, Stig Petersen, et al · 2015
Cited alongside, same era.
Quantilizers: A safer alternative to maximizers for limited optimization
Jessica Taylor · 2016
Cited alongside, same era.
How useful is quantilization for mitigating specification gaming?
Ryan Carey · 2019
Cited alongside, same era.
Human compatible: Artificial intelligence and the problem of control
Stuart Russell · 2019
Later among the works it cites.
Is power-seeking AI an existential risk?, 2021
Joe Carlsmith · 2021
Later among the works it cites.
First return, then explore
Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth O Stanley, and Jeff Clune · 2021
Later among the works it cites.
Optimal policies tend to seek power
Alexander Matt Turner, Logan Smith, Rohin Shah, Andrew Critch, and Prasad Tadepalli · 2021
Later among the works it cites.
Reward is not the optimization target, 2022
Alexander Matt Turner · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…