Fetching the paper…
Reading the bibliography…
Reward function specification can be difficult.
Constrained Markov decision processes , volume 7
Eitan Altman · 1999
Earlier work this paper cites.
Turing universality of the game of life
Paul Rendell · 2002
Earlier work this paper cites.
Robust policy computation in reward-uncertain MDPs using nondominated policies
Kevin Regan and Craig Boutilier · 2010
Earlier work this paper cites.
Superintelligence
Nick Bostrom · 2014
Earlier work this paper cites.
The Frame Problem in Artificial Intelligence: Proceedings of the 1987 Workshop
Frank M Brown · 2014
Earlier work this paper cites.
Safe exploration techniques for reinforcement learning–an overview
Martin Pecka and Tomas Svoboda · 2014
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Javier García and Fernando Fernández · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Cited alongside, same era.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Cited alongside, same era.
Safe model-based reinforcement learning with stability guarantees
Felix Berkenkamp, Matteo Turchetta, Angela Schoellig, and Andreas Krause · 2017
Cited alongside, same era.
Jan Leike, Miljan Martic, Victoria Krakovna, Pedro Ortega, Tom Everitt, Andrew Lefrancq, Laurent Orseau, and Shane Legg · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Towards empathic deep Q-learning
Bart Bussmann, Jacqueline Heinerman, and Joel Lehman · 2019
Later among the works it cites.
The continuous Bernoulli: fixing a pervasive error in variational autoencoders
Gabriel Loaiza-Ganem and John P Cunningham · 2019
Later among the works it cites.
Benchmarking safe exploration in deep reinforcement learning
Alex Ray, Joshua Achiam, and Dario Amodei · 2019
Later among the works it cites.
Human compatible: Artificial intelligence and the problem of control
Stuart Russell · 2019
Later among the works it cites.
The implicit preference information in an initial state
Rohin Shah, Dmitrii Krasheninnikov, Jordan Alexander, Pieter Abbeel, and Anca Dragan · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Leave no trace: Learning to reset for safe and autonomous reinforcement learning
Benjamin Eysenbach, Shixiang Gu, Julian Ibarz, and Sergey Levine · 2018
Cited alongside, same era.
Measuring and avoiding side effects using relative reachability
Victoria Krakovna, Laurent Orseau, Miljan Martic, and Shane Legg · 2018
Cited alongside, same era.
Minimax-regret querying on side effects for safe optimality in factored Markov decision processes
Shun Zhang, Edmund H Durfee, and Satinder P Singh · 2018
Cited alongside, same era.
Carroll L Wainwright and Peter Eckersley · 2019
Later among the works it cites.
Specification gaming: the flip side of AI ingenuity, 2020
Victoria Krakovna, Jonathan Uesato, Vladimir Mikulik, Matthew Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, and Shane Legg · 2020
Closest in time.
Conservative agency via attainable utility preservation
Alexander Matt Turner, Dylan Hadfield-Menell, and Prasad Tadepalli · 2020
Closest in time.