2024

Mechanistic Interpretability for AI Safety -- A Review

Bereska, Leonard, Gavves, Efstratios

Understand

Understanding AI systems' inner workings is critical for ensuring value alignment and safety.

  • This review explores mechanistic interpretability: reverse engineering the computational mechanisms and representations learned by neural networks into human-understandable algorithms and concepts to provide a granular, causal understanding.
  • We establish foundational concepts such as features encoding knowledge within neural activations and hypotheses about their representation and computation.
  • We survey methodologies for causally dissecting model behaviors and assess the relevance of mechanistic interpretability to AI safety.

Reading the bibliography…