Fetching the paper…

Codebook Features: Sparse and Discrete Interpretability for Neural Networks · Around