Fetching the paper…

Automated Interpretability Metrics Do Not Distinguish Trained and Random Transformers · Around