Fetching the paper…

SafeSwitch: Steering Unsafe LLM Behavior via Internal Activation Signals · Around