Fetching the paper…
Reading the bibliography…
Can humans get arbitrarily capable reinforcement learning (RL) agents to do their bidding? Or will sufficiently capable RL agents always find ways to bypass their intended objectives by shortcutting their reward signal? This question impacts how far RL can be scaled, and whether alternative paradigms must be developed in order to build safe artificial general intelligence.
Abram Demski and Scott Garrabrant · 1902
Earlier work this paper cites.
Tom Everitt, Pedro. Ortega, Elizabeth Barnes and Shane Legg · 1902
Earlier work this paper cites.
“Preferences implicit in the state of the world”
Rohin Shah et al · 1902
Earlier work this paper cites.
“Conservative agency via Attainable Utility Preservation”
Alexander Turner, Dylan Hadfield-Menell and Prasad Tadepalli · 1902
Earlier work this paper cites.
“Risks from Learned Optimization in Advanced Machine Learning Systems”, 2019
Evan Hubinger et al · 1906
Earlier work this paper cites.
“Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model”, 2019
Julian Schrittwieser et al · 1911
Earlier work this paper cites.
“Learning human objectives by evaluating hypothetical behavior”
Siddharth Reddy et al · 1912
Earlier work this paper cites.
“Positive Reinforcement Produced by Electrical Stimulation of Septal Area and other Regions of Rat Brain.”
James Olds and Peter Milner · 1954
Earlier work this paper cites.
“Complete identification methods for the causal hierarchy”
Ilya Shpitser and Judea Pearl · 1979
Earlier work this paper cites.
“Influence Diagrams”
Ronald Howard and James Matheson · 1984
Earlier work this paper cites.
“Compulsive thalamic self-stimulation: a case with metabolic, electrophysiologic and behavioral correlates”
Russell Portenoy et al · 1986
Earlier work this paper cites.
“Probabilistic Evaluation of Counterfactual Queries”
Alexander Balke and Judea Pearl · 1994
Earlier work this paper cites.
“Planning and acting in partially observable stochastic domains”
Leslie Kaelbling, Michael. Littman and Anthony. Cassandra · 1998
Earlier work this paper cites.
“Algorithms for inverse reinforcement learning”
Andrew Ng and Stuart Russell · 2000
Earlier work this paper cites.
“Artificial Intelligence, Values and Alignment”
Iason Gabriel · 2001
Earlier work this paper cites.
“Representing and Solving Decision Problems with Limited Information”
Steffen. Lauritzen and Dennis Nilsson · 2001
Earlier work this paper cites.
“Reward-rational (implicit) choice: A unifying formalism for reward learning”, 2020
Hong Jeon, Smitha Milli and Anca. Dragan · 2002
Earlier work this paper cites.
“Multi-agent influence diagrams for representing and solving games”
Daphne Koller and Brian Milch · 2003
Earlier work this paper cites.
“Pitfalls in learning a reward function online”
Stuart Armstrong, Laurent Orseau, Jan Leike and Shane Legg · 2004
Earlier work this paper cites.
“Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems”, 2020
Sergey Levine, Aviral Kumar, George Tucker and Justin Fu · 2005
Earlier work this paper cites.
“Gödel Machines: Self-Referential Universal Problem Solvers Making Provably Optimal Self-Improvements”
Jürgen Schmidhuber · 2007
Earlier work this paper cites.
“What counterfactuals can be tested”
Ilya Shpitser and Judea Pearl · 2007
Earlier work this paper cites.
“From Optimizing Engagement to Measuring Value”
Smitha Milli, Luca Belli and Moritz Hardt · 2008
Cited alongside, same era.
“The Basic AI Drives”
Stephen Omohundro · 2008
Cited alongside, same era.
“Erotic self-stimulation and brain implants”
Vaughanbell · 2008
Cited alongside, same era.
“Hard Takeoff”
Eliezer Yudkowsky · 2008
Cited alongside, same era.
“Interactively shaping agents via human reinforcement”
W. Knox and Peter Stone · 2009
Cited alongside, same era.
“Causality: Models, Reasoning, and Inference”
Judea Pearl · 2009
Cited alongside, same era.
“Learning what to Value”
“Safely interruptible agents”
Laurent Orseau and Stuart Armstrong · 2016
Later among the works it cites.
“‘Indifference’ methods for managing agent rewards”, 2017, pp. 1–16
Stuart Armstrong and Xavier O’Rourke · 2017
Later among the works it cites.
“Deep reinforcement learning from human preferences”
Paul Christiano et al · 2017
Later among the works it cites.
“From Bacteria to Bach and Back: The Evolution of Minds”
Daniel Dennett · 2017
Later among the works it cites.
“Reinforcement Learning with Corrupted Reward Signal”
Tom Everitt et al · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel Dewey · 2011
Cited alongside, same era.
“REALab: An Embedded Perspective on Tampering”, 2020
Ramana Kumar et al · 2011
Cited alongside, same era.
“Self-modification and mortality in artificial agents”
Laurent Orseau and Mark Ring · 2011
Cited alongside, same era.
“Delusion, Survival, and Intelligent Agents”
Mark Ring and Laurent Orseau · 2011
Cited alongside, same era.
“Avoiding Tampering Incentives in Deep RL via Decoupled Approval”, 2020
Jonathan Uesato et al · 2011
Cited alongside, same era.
“Model-based Utility Functions”
Bill Hibbard · 2012
Cited alongside, same era.
Dylan Hadfield-Menell et al · 2017
Later among the works it cites.
Jan Leike et al · 2017
Later among the works it cites.
“Should robots be obedient?”
Smitha Milli, Dylan Hadfield-Menell, Anca Dragan and Stuart Russell · 2017
Later among the works it cites.
“Incorrigibility in the CIRL Framework”
Ryan Carey · 2018
Later among the works it cites.
“Supervising strong learners by amplifying weak experts”, 2018
Paul Christiano, Buck Shlegeris and Dario Amodei · 2018
Later among the works it cites.
“Towards Safe Artificial General Intelligence”, 2018
Tom Everitt · 2018
Later among the works it cites.
“AGI Safety Literature Review”
Tom Everitt, Gary Lea and Marcus Hutter · 2018
Later among the works it cites.
Joel Lehman et al · 2018
Later among the works it cites.
“Scalable agent alignment via reward modeling: a research direction”, 2018
Jan Leike et al · 2018
Later among the works it cites.
“Reinforcement Learning: An Introduction”
Richard Sutton and Andrew Barto · 2018
Later among the works it cites.
“Stuart J. Russell on Filter Bubbles and the Future of Artificial Intelligence”
Stuart Russell · 2019
Closest in time.
“Choice Set Misspecification in Reward Inference”
Rachel Freedman, Rohin Shah and Anca Dragan · 2020
Closest in time.
“Specification gaming: the flip side of AI ingenuity”
Victoria Krakovna et al · 2020
Closest in time.
“Agent Incentives: A Causal Perspective”
Tom Everitt et al · 2021
Closest in time.
“How RL Agents Behave When Their Actions Are Modified”
Eric Langlois and Tom Everitt · 2021
Closest in time.
“Machines Learning Values”
Steve Petersen · 2021
Closest in time.