Fetching the paper…

Improving Neuron-level Interpretability with White-box Language Models · Around