Fetching the paper…
Reading the bibliography…
Power-seeking behavior is a key source of risk from advanced AI, but our theoretical understanding of this phenomenon is relatively limited.
Optimal policies tend to seek power
Alexander Matt Turner, Logan Smith, Rohin Shah, Andrew Critch, and Prasad Tadepalli · 1912
Earlier work this paper cites.
Is power-seeking AI an existential risk?
Joseph Carlsmith · 2022
Earlier work this paper cites.
Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover
Ajeya Cotra · 2022
Earlier work this paper cites.
Goal misgeneralization in deep reinforcement learning
Lauro Langosco, Jack Koch, Lee Sharkey, Jacob Pfau, Laurent Orseau, and David Krueger · 2022
Cited alongside, same era.
The alignment problem from a deep learning perspective
Richard Ngo · 2022
Cited alongside, same era.
Goal misgeneralization: Why correct specifications aren’t enough for correct goals
Rohin Shah, Vikrant Varma, Ramana Kumar, Mary Phuong, Victoria Krakovna, Jonathan Uesato, and Zac Kenton · 2022
Cited alongside, same era.
Reward is not the optimization target
Alexander Matt Turner · 2022
Later among the works it cites.
Parametrically retargetable decision-makers tend to seek power
Alexander Matt Turner and Prasad Tadepalli · 2022
Later among the works it cites.
Definitions of “objective" should be probable and predictive
Rohin Shah · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…