Fetching the paper…
Reading the bibliography…
We present a suite of reinforcement learning environments illustrating various safety properties of intelligent agents.
Iterative Solution of Games by Fictitious Play
George W. Brown · 1951
Earlier work this paper cites.
Game Theory
Drew Fudenberg and Jean Tirole · 1991
Earlier work this paper cites.
Q-learning
Christopher Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald Williams · 1992
Earlier work this paper cites.
The first law of robotics (a call to arms)
Daniel Weld and Oren Etzioni · 1994
Earlier work this paper cites.
Optimal Control: Basics and Beyond
Peter Whittle · 1996
Earlier work this paper cites.
Essentials of Robust Control
Kemin Zhou and John C Doyle · 1997
Earlier work this paper cites.
The MNIST database of handwritten digits, 1998
Yann LeCun · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard Sutton and Andrew Barto · 1998
Earlier work this paper cites.
Risk-sensitive and minimax control of discrete-time, finite-state Markov decision processes
Stefano Coraluppi and Steven Marcus · 1999
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Ng and Stuart Russell · 2000
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert Schapire · 2002
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Ng · 2004
Earlier work this paper cites.
Brown’s original fictitious play
Ulrich Berger · 2007
Earlier work this paper cites.
Safe exploration for reinforcement learning
Alexander Hans, Daniel Schneegaß, Anton Maximilian Schäfer, and Steffen Udluft · 2008
Earlier work this paper cites.
The basic AI drives
Stephen Omohundro · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Dataset Shift in Machine Learning
Joaquin Quiñonero Candela, Masashi Sugiyama, Anton Schwaighofer, and Neil Lawrence · 2009
Earlier work this paper cites.
Autonomous helicopter aerobatics through apprenticeship learning
Pieter Abbeel, Adam Coates, and Andrew Ng · 2010
Earlier work this paper cites.
Self-modification and mortality in artificial agents
Laurent Orseau and Mark Ring · 2011
Earlier work this paper cites.
Delusion, survival, and intelligent agents
Mark Ring and Laurent Orseau · 2011
Earlier work this paper cites.
APRIL: Active preference learning-based reinforcement learning
Riad Akrour, Marc Schoenauer, and Michèle Sebag · 2012
Earlier work this paper cites.
The best of both worlds: stochastic and adversarial bandits
Sebastian Bubeck and Alexander Slivkins · 2012
Earlier work this paper cites.
Model-based utility functions
Bill Hibbard · 2012
Earlier work this paper cites.
Safe exploration in Markov decision processes
Teodor Mihai Moldovan and Pieter Abbeel · 2012
Earlier work this paper cites.
Space-time embedded intelligence
Laurent Orseau and Mark Ring · 2012
Earlier work this paper cites.
Lecture 6.5-RMSProp: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Risk sensitive path integral control
Bart van den Broek, Wim Wiegerinck, and Hilbert Kappen · 2012
Earlier work this paper cites.
A Bayesian approach for policy learning from trajectory preference queries
Aaron Wilson, Alan Fern, and Prasad Tadepalli · 2012
Cited alongside, same era.
The Arcade Learning Environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Cited alongside, same era.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Cited alongside, same era.
Superintelligence: Paths, Dangers, Strategies
Nick Bostrom · 2014
Cited alongside, same era.
Explaining and harnessing adversarial examples
Ian Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Safely interruptible agents
Laurent Orseau and Stuart Armstrong · 2016
Later among the works it cites.
Using stories to teach human values to artificial agents
Mark O Riedl and Brent Harrison · 2016
Later among the works it cites.
Should we fear supersmart robots?
Stuart Russell · 2016
Later among the works it cites.
Towards verified artificial intelligence
Sanjit A Seshia, Dorsa Sadigh, and S Shankar Sastry · 2016
Later among the works it cites.
Quantilizers: A safer alternative to maximizers for limited optimization
Jessica Taylor · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Transcending complacency on superintelligent machines
Stephen Hawking, Max Tegmark, Stuart Russell, and Frank Wilczek · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Safe exploration techniques for reinforcement learning — an overview
Martin Pecka and Tomas Svoboda · 2014
Cited alongside, same era.
Empowerment—an introduction
Christoph Salge, Cornelius Glackin, and Daniel Polani · 2014
Cited alongside, same era.
One practical algorithm for both stochastic and adversarial bandits
Yevgeny Seldin and Alexander Slivkins · 2014
Cited alongside, same era.
Aligning superintelligence with human interests: A technical research agenda
Nate Soares and Benja Fallenstein · 2014
Cited alongside, same era.
A unified framework for risk-sensitie Markov control processes
Shen Yun, Wilhelm Stannat, and Klaus Obermayer · 2014
Cited alongside, same era.
Later among the works it cites.
Alignment for advanced machine learning systems
Jessica Taylor, Eliezer Yudkowsky, Patrick LaVictoire, and Andrew Critch · 2016
Later among the works it cites.
Safe exploration in finite Markov decision processes with Gaussian processes
Matteo Turchetta, Felix Berkenkamp, and Andreas Krause · 2016
Later among the works it cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Closest in time.
AI toy control problem, 2017
Stuart Armstrong · 2017
Closest in time.
Low impact artificial intelligences
Stuart Armstrong and Benjamin Levinstein · 2017
Closest in time.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Closest in time.
Agent coordination and potential risks: Meaningful environments for evaluating multiagent systems
Nader Chmait, David L Dowe, David G Green, and Yuan-Fang Li · 2017
Closest in time.
Deep reinforcement learning from human preferences
Paul Christiano, Jan Leike, Tom B Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Closest in time.
A roadmap for a rigorous science of interpretability
Finale Doshi-Velez and Been Kim · 2017
Closest in time.
Reinforcement learning with corrupted reward signal
Tom Everitt, Victoria Krakovna, Laurent Orseau, Marcus Hutter, and Shane Legg · 2017
Closest in time.
As emissions scandal widens, diesel’s future looks shaky in Europe, July 2017
Jack Ewing · 2017
Closest in time.
Noisy networks for exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Ian Osband, Alex Graves, Vlad Mnih, Remi Munos, Demis Hassabis, Olivier Pietquin, Charles Blundell, and Shane Legg · 2017
Closest in time.
Inverse reward design
Dylan Hadfield-Menell, Smitha Milli, Stuart Russell, Pieter Abbeel, and Anca D Dragan · 2017
Closest in time.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2017
Closest in time.
Learning from demonstrations for real world reinforcement learning
Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Andrew Sendonaris, Gabriel Dulac-Arnold, Ian Osband, John Agapiou, Joel Z Leibo, and Audrunas Gruslys · 2017
Closest in time.
Towards proving the adversarial robustness of deep neural networks
Guy Katz, Clark Barrett, David Dill, Kyle Julian, and Mykel Kochenderfer · 2017
Closest in time.
Interactive learning from policy-dependent human feedback
James MacGlashan, Mark K Ho, Robert Loftin, Bei Peng, David Roberts, Matthew E Taylor, and Michael L Littman · 2017
Closest in time.
Should robots be obedient?
Smitha Milli, Dylan Hadfield-Menell, Anca Dragan, and Stuart Russell · 2017
Closest in time.
Enter the matrix: A virtual world approach to safely interruptable autonomous systems
Mark O Riedl and Brent Harrison · 2017
Closest in time.
RAIL: Risk-averse imitation learning
Anirban Santara, Abhishek Naik, Balaraman Ravindran, Dipankar Das, Dheevatsa Mudigere, Sasikanth Avancha, and Bharat Kaul · 2017
Closest in time.
Trial without error: Towards safe reinforcement learning via human intervention
William Saunders, Girish Sastry, Andreas Stuhlmueller, and Owain Evans · 2017
Closest in time.
The pycolab game engine, 2017
Thomas Stepleton · 2017
Closest in time.
A Berkeley view of systems challenges for AI
Ion Stoica, Dawn Song, Raluca Ada Popa, David A Patterson, Michael W Mahoney, Randy H Katz, Anthony D Joseph, Michael Jordan, Joseph M Hellerstein, Joseph Gonzalez, et al · 2017
Closest in time.