Fetching the paper…
Reading the bibliography…
Explainable reinforcement learning (XRL) is an emerging subfield of explainable machine learning that has attracted considerable attention in recent years.
Apprenticeship learning via inverse reinforcement learning
P Abbeel and AY Ng · 2004
Earlier work this paper cites.
A survey of robot learning from demonstration
BD Argall, S Chernova, et al · 2009
Earlier work this paper cites.
Playing Atari with deep reinforcement learning
V Mnih, K Kavukcuoglu, et al · 2013
Earlier work this paper cites.
Graying the black box: Understanding dqns
T Zahavy, N Ben-Zrihem, and S Mannor · 2016
Earlier work this paper cites.
Explainable artificial intelligence
D Gunning · 2017
Earlier work this paper cites.
Improving robot controller transparency through autonomous policy explanation
B Hayes and JA Shah · 2017
Earlier work this paper cites.
Particle swarm optimization for generating interpretable fuzzy reinforcement learning policies
D Hein, A Hentschel, et al · 2017
Earlier work this paper cites.
Highlights: Summarizing agent behavior to people
D Amir and O Amir · 2018
Earlier work this paper cites.
Multimodal machine learning: A survey and taxonomy
T Baltrušaitis, C Ahuja, and LP Morency · 2018
Earlier work this paper cites.
Verifiable reinforcement learning via policy extraction
O Bastani, Y Pu, and A Solar-Lezama · 2018
Earlier work this paper cites.
Deep reinforcement learning monitor for snapshot recording
G Dao, I Mishra, and M Lee · 2018
Earlier work this paper cites.
Rationalization: A neural machine translation approach to generating natural language explanations
U Ehsan, B Harrison, et al · 2018
Earlier work this paper cites.
Unsupervised video object segmentation for deep reinforcement learning
V Goel, J Weng, and P Poupart · 2018
Earlier work this paper cites.
Visualizing and understanding atari agents
S Greydanus, A Koul, et al · 2018
Earlier work this paper cites.
Establishing appropriate trust via critical states
SH Huang, K Bhatia, et al · 2018
Earlier work this paper cites.
Transparency and explanation in deep reinforcement learning neural networks
R Iyer, Y Li, et al · 2018
Earlier work this paper cites.
The mythos of model interpretability
ZC Lipton · 2018
Earlier work this paper cites.
Toward interpretable deep reinforcement learning with linear model u-trees
G Liu, O Schulte, et al · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
RS Sutton and AG Barto · 2018
Earlier work this paper cites.
Programmatically interpretable reinforcement learning
A Verma, V Murali, et al · 2018
Earlier work this paper cites.
Visual rationalizations in deep reinforcement learning for Atari games
L Weitkamp, E van der Pol, and Z Akata · 2018
Earlier work this paper cites.
Explaining reinforcement learning to mere mortals: An empirical study
A Anderson, J Dodge, et al · 2019
Earlier work this paper cites.
Towards better interpretability in deep q-networks
RM Annasamy and K Sycara · 2019
Earlier work this paper cites.
Dot-to-dot: Explainable hierarchical reinforcement learning for robotic manipulation
B Beyret, A Shafti, and AA Faisal · 2019
Earlier work this paper cites.
Memory-based explainable reinforcement learning
F Cruz, R Dazeley, and P Vamplew · 2019
Cited alongside, same era.
Interpretable policies for reinforcement learning by genetic programming
D Hein, S Udluft, and TA Runkler · 2019
Cited alongside, same era.
Enhancing explainability of deep reinforcement learning through selective layer-wise relevance propagation
T Huber, D Schiller, and E André · 2019
Cited alongside, same era.
Policy extraction via online q-value distillation
A Jhunjhunwala · 2019
Cited alongside, same era.
Explainable reinforcement learning via reward decomposition
Z Juozapaitis, A Koul, et al · 2019
Cited alongside, same era.
Learning finite state representations of recurrent policy networks
A Koul, S Greydanus, and A Fern · 2019
Cited alongside, same era.
Tldr: Policy summarization for factored ssp problems using temporal abstractions
S Sreedharan, S Srivastava, and S Kambhampati · 2020
Later among the works it cites.
Neuroevolution of self-interpretable agents
Y Tang, D Nguyen, and D Ha · 2020
Later among the works it cites.
What did you think would happen? explaining agent behaviour through intended outcomes
H Yau, C Russell, and S Hadfield · 2020
Later among the works it cites.
Interpretable policy derivation for reinforcement learning based on evolutionary feature synthesis
H Zhang, A Zhou, and X Lin · 2020
Later among the works it cites.
Tripletree: A versatile interpretable representation of black box agents and their environments
T Bewley and J Lawry · 2021
Later among the works it cites.
Learning “what-if” explanations for sequential decision-making
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Exploring computational user models for agent policy summarization
I Lage, D Lifschitz, et al · 2019
Cited alongside, same era.
Interpretable Machine Learning
C Molnar · 2019
Cited alongside, same era.
Towards interpretable reinforcement learning using attention augmented agents
A Mott, D Zoran, et al · 2019
Cited alongside, same era.
Definitions, methods, and applications in interpretable machine learning
WJ Murdoch, C Singh, et al · 2019
Cited alongside, same era.
Generation of policy-level explanations for reinforcement learning
N Topin and M Veloso · 2019
Cited alongside, same era.
Verbal explanations for deep reinforcement learning neural networks with attention on extracted features
X Wang, Y Schengcheng, et al · 2019
Cited alongside, same era.
I Bica, D Jarrett, et al · 2021
Later among the works it cites.
Explainable robotic systems: Understanding goal-driven actions in a reinforcement learning scenario
F Cruz, R Dazeley, and P Vamplew · 2021
Later among the works it cites.
Re-understanding finite-state representations of recurrent policy networks
MH Danesh, A Koul, et al · 2021
Later among the works it cites.
From “no clear winner” to an effective explainable artificial intelligence process: An empirical journey
J Dodge, A Anderson, et al · 2021
Later among the works it cites.
Edge: Explaining deep reinforcement learning policies
W Guo, X Wu, et al · 2021
Later among the works it cites.
Deepsynth: Automata synthesis for automatic task segmentation in deep reinforcement learning
M Hasanbeig, NY Jeppu, et al · 2021
Later among the works it cites.
Explainability in deep reinforcement learning
A Heuillet, F Couthouis, and N Díaz-Rodríguez · 2021
Later among the works it cites.
Visual explanation using attention mechanism in actor-critic-based deep reinforcement learning
H Itaya, T Hirakawa, et al · 2021
Later among the works it cites.
Discovering symbolic policies with deep reinforcement learning
M Landajuela, BK Petersen, et al · 2021
Later among the works it cites.
Contrastive explanations for reinforcement learning via embedded self predictions
Z Lin, KH Lam, and A Fern · 2021
Later among the works it cites.
Counterfactual state explanations for reinforcement learning agents via generative deep learning
ML Olson, R Khanna, et al · 2021
Later among the works it cites.
The MineRL basalt competition on learning from human feedback
R Shah, C Wild, et al · 2021
Later among the works it cites.
Iterative bounding mdps: Learning interpretable policies via non-interpretable methods
N Topin, S Milani, et al · 2021
Later among the works it cites.
Causeoccam: Learning interpretable abstract representations in reinforcement learning environments via model sparsity
S Volodin · 2021
Later among the works it cites.
Explainable ai and reinforcement learning—a systematic review of current approaches and trends
L Wells and T Bednarz · 2021
Later among the works it cites.
Assessing explainability in reinforcement learning
AE Zelvelder, M Westberg, and K Främling · 2021
Later among the works it cites.
Off-policy differentiable logic reinforcement learning
L Zhang, X Li, et al · 2021
Later among the works it cites.
Learning to discover task-relevant features for interpretable reinforcement learning
Q Zhang, X Ma, et al · 2021
Later among the works it cites.