Fetching the paper…
Reading the bibliography…
Many works in explainable AI have focused on explaining black-box classification models.
Dynamic thick restarting of the davidson, and the implicitly restarted arnoldi methods
A. Stathopoulos, Y. Saad, and K. Wu · 1997
Earlier work this paper cites.
Arpack users guide: Solution of large-scale eigenvalue problems with implicitly restarted arnoldi methods
R. B. Lehoucq, D. C. Sorensen, and C. Yang · 1998
Earlier work this paper cites.
Normalized cuts and image segmentation
Jianbo Shi and Jitendra Malik · 2000
Earlier work this paper cites.
Q-cut - dynamic discovery of sub-goals in reinforcement learning
Ishai Menache, Shie Mannor, and Nahum Shimkin · 2002
Earlier work this paper cites.
Hierarchical interpretations for neural network predictions
O. Simsek and A. G. Barto · 2004
Earlier work this paper cites.
A tutorial of spectral clustering
U. von Luxburg · 2007
Earlier work this paper cites.
Hierarchically organized behavior and its neural foundations: a reinforcement learning perspective
MIchael M. Botvinick, Yael Niv, and Andrew C. Barto · 2008
Earlier work this paper cites.
Generating explanations based on markov decision processes
Francisco Elizalde, Enrique Sucar, Julieta Noguez, and Alberto Reyes · 2009
Earlier work this paper cites.
Minimal sufficient explanations for factored markov decision processes
Omar Zia Khan, Pascal Poupart, and James P. Black · 2009
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, Silver Kavukcuoglu, K., and D. et al · 2015
Earlier work this paper cites.
Personalized ad recommendation systems for life-time value optimization with guarantees
G. Theocharous, P. S. Thomas, and M. Ghavamzadeh · 2015
Earlier work this paper cites.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
The lrp toolbox for artificial neural networks
Sebastian Lapuschkin, Alexander Binder, Grégoire Montavon, Klaus-Robert Müller, and Wojciech Samek · 2016
Earlier work this paper cites.
State of the art control of atari games using shallow reinforcement learning
Yitao Liang, Marlos C. Machado, Erik Talvitie, and Michael Bowling · 2016
Earlier work this paper cites.
The mythos of model interpretability
Zachary C Lipton · 2016
Earlier work this paper cites.
“Why should I trust you?”: Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, Maddison Huang, A., and C. et al · 2016
Earlier work this paper cites.
Graying the black gox: Understanding dqns
Tom Zahavy, Nir Ben Zrihem, and Shie Mannor · 2016
Cited alongside, same era.
TIP: Typifying the interpretability of procedures
Amit Dhurandhar, Vijay Iyengar, Ronny Luss, and Karthikeyan Shanmugam · 2017
Cited alongside, same era.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim · 2017
Cited alongside, same era.
Explainable artificial intelligence (xai)
David Gunning · 2017
Cited alongside, same era.
A unified approach to interpreting model predictions
Scott M. Lundberg and Su-In Lee · 2017
Cited alongside, same era.
Highlights: Summarizing agent behavior to people
Dan Amir and Ofra Amir · 2018
Exploring computational user models for agent policy summarization
Isaac Lage, Daphna Lifschitz, Finale Doshi-Velez, and Ofra Amir · 2019
Later among the works it cites.
Towards interpretable reinforcement learning using attention augmented agents
Alexander Mott, Daniel Zoran, Mike Chrzanowski, Daan Wierstra, and Danilo Jimenez Rezende · 2019
Later among the works it cites.
Learning from trajectories via subgoal discovery
Sujoy Paul, Jeroen Vanbaar, and Amit Roy-Chowdhury · 2019
Later among the works it cites.
Successor options: An option discovery framework for reinforcement learning
Rahul Ramesh, Manan Tomar, and Balaraman Ravindran · 2019
Later among the works it cites.
Generation of policy-level explanations for reinforcement learning
Nicholay Topin and Manuela Veloso · 2019
Later among the works it cites.
Reinforcement learning interpretation methods: A survey
Alnour Alharin, Thanh-Nam Doan, and Mina Sartipi · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Verifiable reinforcement learning via policy extraction
Osbert Bastani, Yewen Pu, and Armando Solar-Lezama · 2018
Cited alongside, same era.
Explanations based on the missing: Towards contrastive explanations with pertinent negatives
Amit Dhurandhar, Pin-Yu Chen, Ronny Luss, Chun-Chen Tu, Paishun Ting, Karthikeyan Shanmugam, and Payel Das · 2018
Cited alongside, same era.
Visualizing and understanding atari agents
Samuel Greydanus, Anurag Koul, Jonathan Dodge, and Alan Fern · 2018
Cited alongside, same era.
Establishing appropriate trust via critical states
Sandy H. Huang, Kush Bhatia, Pieter Abbeel, and Anca D. Dragan · 2018
Cited alongside, same era.
Eigenoption discovery through the deep successor representation
Marlos C. Machado, Clemens Rosenbaum, Xiaoxiao Guo, Miao Liu, Gerald Tesauro, and Murray Campbell · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2018
Cited alongside, same era.
Later among the works it cites.
Synthesizing programmatic policies that inductively generalize
Jeevana Priya Inala, Osbert Bastani, Zenna Tavares, and Armando Solar-Lezama · 2020
Later among the works it cites.
Explainable reinforcement learning through a causal lens
Prashan Madumal, Tim Miller, Liz Sonenberg, and Frank Vetere · 2020
Later among the works it cites.
Finding and visualizing weaknesses of deep reinforcement learning agents
Christian Rupprecht, Cyril Ibrahim, and Christopher J. Pal · 2020
Later among the works it cites.
Tldr: Policy summarization for factored ssp problems using temporal abstractions
S. Sreedharan, S. Srivastava, and S. Kambhampati · 2020
Later among the works it cites.
What did you think would happen? explaining agent behaviour through intended outcomes
Herman Yau, Chris Russell, and Simon Hadfield · 2020
Later among the works it cites.
A survey on the explainability of supervised machine learning
N. Burkart and M. C. Huber · 2021
Later among the works it cites.
Re-understanding finite state representations of recurrent policy networks
Mohamad H. Danesh, Anurag Koul, Alan Fern, and Saeed Khorram · 2021
Later among the works it cites.
Local and global explanations of agent behavior: Integrating strategy summaries with saliency maps
Tobias Huber, Katharina Weitz, Elisabeth André, and Ofra Amir · 2021
Later among the works it cites.
Learning tree interpretation from object representation for deep reinforcement learning
G. Liu, X. Sun, O. Schulte, and P. Poupart · 2021
Later among the works it cites.
Counterfactual state explanations for reinforcement learning agents via generative deep learning
Matthew L. Olson, Roli Khanna, Lawrence Neal, Fuxin Li, and Weng-Keen Wong · 2021
Later among the works it cites.
Learn from the best: Alphazero, 2021
Daniel Rensch · 2021
Later among the works it cites.