2023

Faithful Explanations of Black-box NLP Models Using LLM-generated Counterfactuals

Gat, Yair, Calderon, Nitay, Feder, Amir et al.

Understand

Causal explanations of the predictions of NLP systems are essential to ensure safety and establish trust.

  • Yet, existing methods often fall short of explaining model predictions effectively or efficiently and are often model-specific.
  • In this paper, we address model-agnostic explanations, proposing two approaches for counterfactual (CF) approximation.
  • The first approach is CF generation, where a large language model (LLM) is prompted to change a specific text concept while keeping confounding concepts unchanged.

Reading the bibliography…