Fetching the paper…

Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models · Around