2024

A Multimodal Automated Interpretability Agent

Shaham, Tamar Rott, Schwettmann, Sarah, Wang, Franklin et al.

Understand

This paper describes MAIA, a Multimodal Automated Interpretability Agent.

  • MAIA is a system that uses neural models to automate neural model understanding tasks like feature interpretation and failure mode discovery.
  • It equips a pre-trained vision-language model with a set of tools that support iterative experimentation on subcomponents of other models to explain their behavior.
  • These include tools commonly used by human interpretability researchers: for synthesizing and editing inputs, computing maximally activating exemplars from real-world datasets, and summarizing and describing experimental results.

Reading the bibliography…