Fetching the paper…
Reading the bibliography…
Reward functions are easy to misspecify; although designers can make corrections after observing mistakes, an agent pursuing a misspecified reward function can irreversibly change the state of its environment.
Q-learning
Christopher Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Constrained Markov decision processes
Eitan Altman · 1999
Earlier work this paper cites.
Robust policy computation in reward-uncertain MDPs using nondominated policies
Kevin Regan and Craig Boutilier · 2010
Earlier work this paper cites.
Safe exploration in Markov decision processes
Teodor Mihai Moldovan and Pieter Abbeel · 2012
Earlier work this paper cites.
Safe exploration techniques for reinforcement learning–an overview
Martin Pecka and Tomas Svoboda · 2014
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Javier García and Fernando Fernández · 2015
Earlier work this paper cites.
Variational information maximisation for intrinsically motivated reinforcement learning
Shakir Mohamed and Danilo Jimenez Rezende · 2015
Earlier work this paper cites.
Corrigibility
Nate Soares, Benja Fallenstein, Stuart Armstrong, and Eliezer Yudkowsky · 2015
Earlier work this paper cites.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Cooperative inverse reinforcement learning
Dylan Hadfield-Menell, Stuart Russell, Pieter Abbeel, and Anca Dragan · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Low impact artificial intelligences
Stuart Armstrong and Benjamin Levinstein · 2017
Cited alongside, same era.
Safe model-based reinforcement learning with stability guarantees
Felix Berkenkamp, Matteo Turchetta, Angela Schoellig, and Andreas Krause · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Incorrigibility in the CIRL framework
Ryan Carey · 2018
Later among the works it cites.
A lyapunov-based approach to safe reinforcement learning
Yinlam Chow, Ofir Nachum, Edgar Duenez-Guzman, and Mohammad Ghavamzadeh · 2018
Later among the works it cites.
Leave no trace: Learning to reset for safe and autonomous reinforcement learning
Benjamin Eysenbach, Shixiang Gu, Julian Ibarz, and Sergey Levine · 2018
Later among the works it cites.
Measuring and avoiding side effects using relative reachability
Victoria Krakovna, Laurent Orseau, Miljan Martic, and Shane Legg · 2018
Later among the works it cites.
Preventing side-effects in gridworlds, 2018
Gavin Leech, Karol Kubicki, Jessica Cooper, and Tom McGrath · 2018
Later among the works it cites.
OpenAI Five
OpenAI · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tom Everitt, Victoria Krakovna, Laurent Orseau, and Shane Legg · 2017
Cited alongside, same era.
The off-switch game
Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell · 2017
Cited alongside, same era.
Inverse reward design
Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, Stuart Russell, and Anca Dragan · 2017
Cited alongside, same era.
Jan Leike, Miljan Martic, Victoria Krakovna, Pedro Ortega, Tom Everitt, Andrew Lefrancq, Laurent Orseau, and Shane Legg · 2017
Cited alongside, same era.
Later among the works it cites.
Trial without error: Towards safe reinforcement learning via human intervention
William Saunders, Girish Sastry, Andreas Stuhlmueller, and Owain Evans · 2018
Later among the works it cites.
Minimax-regret querying on side effects for safe optimality in factored Markov decision processes
Shun Zhang, Edmund H Durfee, and Satinder P Singh · 2018
Later among the works it cites.
The implicit preference information in an initial state
Rohin Shah, Dmitrii Krasheninnikov, Jordan Alexander, Pieter Abbeel, and Anca Dragan · 2019
Closest in time.