Fetching the paper…
Reading the bibliography…
One of the most prominent methods for explaining the behavior of Deep Reinforcement Learning (DRL) agents is the generation of saliency maps that show how much each pixel attributed to the agents' decision.
The arcade learning environment: An evaluation platform for general agents
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
”Why should i trust you?” Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot, and Nando de Freitas · 2016
Earlier work this paper cites.
Graying the black box: Understanding dqns
Tom Zahavy, Nir Ben-Zrihem, and Shie Mannor · 2016
Earlier work this paper cites.
Openai baselines
Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, Yuhuai Wu, and Peter Zhokhov · 2017
Earlier work this paper cites.
Evaluating the visualization of what a deep neural network has learned
Wojciech Samek, Alexander Binder, Grégoire Montavon, Sebastian Lapuschkin, and Klaus-Robert Müller · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Earlier work this paper cites.
HIGHLIGHTS: summarizing agent behavior to people
Dan Amir and Ofra Amir · 2018
Cited alongside, same era.
Towards better understanding of gradient-based attribution methods for deep neural networks
Marco Ancona, Enea Ceolini, Cengiz Öztireli, and Markus Gross · 2018
Cited alongside, same era.
Visualizing and understanding atari agents
Samuel Greydanus, Anurag Koul, Jonathan Dodge, and Alan Fern · 2018
Cited alongside, same era.
Learning how to explain neural networks: Patternnet and patternattribution
Pieter-Jan Kindermans, Kristof T. Schütt, Maximilian Alber, Klaus-Robert Müller, Dumitru Erhan, Been Kim, and Sven Dähne · 2018
Cited alongside, same era.
RISE: randomized input sampling for explanation of black-box models
Vitali Petsiuk, Abir Das, and Kate Saenko · 2018
Cited alongside, same era.
Explaining reinforcement learning to mere mortals: An empirical study
Andrew Anderson, Jonathan Dodge, Amrita Sadarangani, Zoe Juozapaitis, Evan Newman, Jed Irvine, Souti Chattopadhyay, Alan Fern, and Margaret Burnett · 2019
A multidisciplinary survey and framework for design and evaluation of explainable ai systems, 2020
Sina Mohseni, Niloofar Zarei, and Eric D. Ragan · 2020
Later among the works it cites.
Explain your move: Understanding agent actions using specific and relevant feature attribution
Nikaash Puri, Sukriti Verma, Piyush Gupta, Dhruv Kayastha, Shripad Deshmukh, Balaji Krishnamurthy, and Sameer Singh · 2020
Later among the works it cites.
Restricting the flow: Information bottlenecks for attribution
Karl Schulz, Leon Sixt, Federico Tombari, and Tim Landgraf · 2020
Later among the works it cites.
When explanations lie: Why many modified BP attributions fail
Leon Sixt, Maximilian Granz, and Tim Landgraf · 2020
Later among the works it cites.
Sanity checks for saliency metrics
Richard Tomsett, Dan Harborne, Supriyo Chakraborty, Prudhvi Gurram, and Alun D. Preece · 2020
Later among the works it cites.
Re-understanding finite-state representations of recurrent policy networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Enhancing explainability of deep reinforcement learning through selective layer-wise relevance propagation
Tobias Huber, Dominik Schiller, and Elisabeth André · 2019
Cited alongside, same era.
Exploratory not explanatory: Counterfactual analysis of saliency maps for deep reinforcement learning
Akanksha Atrey, Kaleigh Clary, and David Jensen · 2020
Cited alongside, same era.
Mohamad H. Danesh, Anurag Koul, Alan Fern, and Saeed Khorram · 2021
Closest in time.
Explainability in deep reinforcement learning
Alexandre Heuillet, Fabien Couthouis, and Natalia Díaz Rodríguez · 2021
Closest in time.
Local and global explanations of agent behavior: Integrating strategy summaries with saliency maps
Tobias Huber, Katharina Weitz, Elisabeth André, and Ofra Amir · 2021
Closest in time.