2018

xGEMs: Generating Examplars to Explain Black-Box Models

Joshi, Shalmali, Koyejo, Oluwasanmi, Kim, Been et al.

Understand

This work proposes xGEMs or manifold guided exemplars, a framework to understand black-box classifier behavior by exploring the landscape of the underlying data manifold as data points cross decision boundaries.

  • To do so, we train an unsupervised implicit generative model -- treated as a proxy to the data manifold.
  • We summarize black-box model behavior quantitatively by perturbing data samples along the manifold.
  • We demonstrate xGEMs' ability to detect and quantify bias in model learning and also for understanding the changes in model behavior as training progresses.

Reading the bibliography…