Fetching the paper…
Reading the bibliography…
We present a general framework for training safe agents whose naive incentives are unsafe.
The principles and applications of decision analysis
Howard, R. A.; and Matheson, J. E. 1984 · 1984
Earlier work this paper cites.
Learning from Demonstration
Schaal, S. 1997 · 1997
Earlier work this paper cites.
Stochastic Dynamic Programming with Factored Representations
Boutilier, C.; Dearden, R.; and Goldszmidt, M. 2000 · 2000
Earlier work this paper cites.
Representing and solving decision problems with limited information
Lauritzen, S. L.; and Nilsson, D. 2001 · 2001
Earlier work this paper cites.
Direct and Indirect Effects
Pearl, J. 2001 · 2001
Earlier work this paper cites.
Identifiability of Path-Specific Effects
Avin, C.; Shpitser, I.; and Pearl, J. 2005 · 2005
Earlier work this paper cites.
Influence Diagrams
Howard, R. A.; and Matheson, J. E. 2005 · 2005
Earlier work this paper cites.
Hidden Incentives for Auto-Induced Distributional Shift
Krueger, D.; Maharaj, T.; and Leike, J. 2020 · 2009
Earlier work this paper cites.
Causality: Models, Reasoning, and Inference
Pearl, J. 2009 · 2009
Earlier work this paper cites.
Alternative Graphical Causal Models and the Identification of Direct Effects
Robins, J. M.; and Richardson, T. S. 2011 · 2011
Cited alongside, same era.
Avoiding Tampering Incentives in Deep RL via Decoupled Approval
Uesato, J.; Kumar, R.; Krakovna, V.; Everitt, T.; Ngo, R.; and Legg, S. 2020 · 2011
Cited alongside, same era.
Counterfactual graphical models for longitudinal mediation analysis with unobserved confounding
Shpitser, I. 2013 · 2013
Cited alongside, same era.
Experimental evidence of massive-scale emotional contagion through social networks
Kramer, A. D. I.; Guillory, J. E.; and Hancock, J. T. 2014 · 2014
Cited alongside, same era.
Concrete Problems in AI Safety
Amodei, D.; Olah, C.; Steinhardt, J.; Christiano, P.; Schulman, J.; and Mané, D. 2016 · 2016
Cited alongside, same era.
Towards Formal Definitions of Blameworthiness, Intention, and Moral Responsibility
Halpern, J. Y.; and Kleiman-Weiner, M. 2018 · 2018
Later among the works it cites.
Scalable agent alignment via reward modeling: a research direction
Leike, J.; Krueger, D.; Everitt, T.; Martic, M.; Maini, V.; and Legg, S. 2018 · 2018
Later among the works it cites.
Estimation of Personalized Effects Asssociated with Causal Pathways
Razieh, N.; Kanki, P.; and Shpitser, I. 2018 · 2018
Later among the works it cites.
Human Compatible: Artificial Intelligence and the Problem of Control
Russell, S. 2019 · 2019
Later among the works it cites.
Pitfalls of Learning a Reward Function Online
Armstrong, S.; Leike, J.; Orseau, L.; and Legg, S. 2020 · 2020
Later among the works it cites.
Avoiding Side Effects by Considering Future Tasks
Krakovna, V.; Orseau, L.; Ngo, R.; Martic, M.; and Legg, S. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Taylor, J. 2016 · 2016
Cited alongside, same era.
Deep Reinforcement Learning from Human Preferences
Christiano, P. F.; Leike, J.; Brown, T. B.; Martic, M.; Legg, S.; and Amodei, D. 2017 · 2017
Cited alongside, same era.
Population Based Training of Neural Networks
Jaderberg, M.; Dalibard, V.; Osindero, S.; Czarnecki, W. M.; Donahue, J.; Razavi, A.; Vinyals, O.; Green, T.; Dunning, I.; Simonyan, K.; Fernando, C.; and Kavukcuoglu, K. 2017 · 2017
Cited alongside, same era.
An analysis of the interaction between intelligent software agents and human users
Burr, C.; Cristianini, N.; and Ladyman, J. 2018 · 2018
Cited alongside, same era.
Agent Incentives: A Causal Perspective
Everitt, T.; Carey, R.; Langlois, E.; Ortega, P. A.; and Legg, S. 2021a
Cited in the paper.
Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective
Everitt, T.; Hutter, M.; Kumar, R.; and Krakovna, V. 2021b
Cited in the paper.
Later among the works it cites.
Avoiding Side Effects in Complex Systems
Turner, A.; Ratzlaff, N.; and Tadepalli, P. 2020 · 2020
Later among the works it cites.
Counterfactual control incentives
Armstrong, S.; and Gorman, R. 2021 · 2021
Later among the works it cites.
Estimating and Penalizing Induced Preference Shifts in Recommender Systems
Carroll, M.; Hadfield-Menell, D.; Russell, S.; and Dragan, A. 2021 · 2021
Later among the works it cites.