2024

How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Zhou, Zhenhong, Yu, Haiyang, Zhang, Xinghua et al.

Understand

Large language models (LLMs) rely on safety alignment to avoid responding to malicious user inputs.

  • Unfortunately, jailbreak can circumvent safety guardrails, resulting in LLMs generating harmful content and raising concerns about LLM safety.
  • Due to language models with intensive parameters often regarded as black boxes, the mechanisms of alignment and jailbreak are challenging to elucidate.
  • In this paper, we employ weak classifiers to explain LLM safety through the intermediate hidden states.

Reading the bibliography…