2019

Explain Your Move: Understanding Agent Actions Using Specific and Relevant Feature Attribution

Puri, Nikaash, Verma, Sukriti, Gupta, Piyush et al.

Understand

As deep reinforcement learning (RL) is applied to more tasks, there is a need to visualize and understand the behavior of learned agents.

  • Saliency maps explain agent behavior by highlighting the features of the input state that are most relevant for the agent in taking an action.
  • Existing perturbation-based approaches to compute saliency often highlight regions of the input that are not relevant to the action taken by the agent.
  • Our proposed approach, SARFA (Specific and Relevant Feature Attribution), generates more focused saliency maps by balancing two aspects (specificity and relevance) that capture different desiderata of saliency.

Reading the bibliography…