Fetching the paper…
Reading the bibliography…
In many real-world tasks, it is not possible to procedurally specify an RL agent's reward function.
Some moral and technical consequences of automation
Norbert Wiener · 1960
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Artificial intelligence as a positive and negative factor in global risk
Eliezer Yudkowsky et al · 2008
Earlier work this paper cites.
Polynomial calculation of the shapley value based on sampling
Javier Castro, Daniel Gómez, and Juan Tejada · 2009
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Cooperative inverse reinforcement learning
Dylan Hadfield-Menell, Stuart Russell, Pieter Abbeel, and Anca Dragan · 2016
Earlier work this paper cites.
"Why should i trust you?" explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje · 2017
Cited alongside, same era.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Cited alongside, same era.
Counterfactual explanations without opening the black box: Automated decisions and the gdpr
Sandra Wachter, Brent Mittelstadt, and Chris Russell · 2017
Cited alongside, same era.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Counterfactual visual explanations
Yash Goyal, Ziyan Wu, Jan Ernst, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Later among the works it cites.
Chris Olah’s views on AGI safety, 2019
Evan Hubinger · 2019
Later among the works it cites.
Towards interpretable reinforcement learning using attention augmented agents
Alex Mott, Daniel Zoran, Mike Chrzanowski, Daan Wierstra, and Danilo J. Rezende · 2019
Later among the works it cites.
Finding and visualizing weaknesses of deep reinforcement learning agents
Christian Rupprecht, Cyril Ibrahim, and Christopher J. Pal · 2019
Later among the works it cites.
Explaining reward functions in markov decision processes
Jacob Russell and Eugene Santos · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Visualizing and understanding Atari agents
Samuel Greydanus, Anurag Koul, Jonathan Dodge, and Alan Fern · 2018
Cited alongside, same era.
Reward learning from human preferences and demonstrations in Atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
Cited alongside, same era.
OpenAI Five
OpenAI · 2018
Cited alongside, same era.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Daniel S. Brown, Wonjoon Goo, Prabhat Nagarajan, and Scott Niekum · 2019
Cited alongside, same era.
Causal confusion in imitation learning
Pim de Haan, Dinesh Jayaraman, and Sergey Levine · 2019
Cited alongside, same era.
Alphastar: Grandmaster level in StarCraft II using multi-agent reinforcement learning
The AlphaStar team · 2019
Later among the works it cites.
Multi-principal assistance games
Arnaud Fickinger, Simon Zhuang, Dylan Hadfield-Menell, and Stuart Russell · 2020
Closest in time.
Quantifying differences in reward functions
Adam Gleave, Michael Dennis, Shane Legg, Stuart Russell, and Jan Leike · 2020
Closest in time.
Reward-rational (implicit) choice: A unifying formalism for reward learning
Hong Jun Jeon, Smitha Milli, and Anca D. Dragan · 2020
Closest in time.
Benefits of assistance over reward learning
Anonymous · 2021
Closest in time.