2024

A safety realignment framework via subspace-oriented model fusion for large language models

Yi, Xin, Zheng, Shunfan, Wang, Linlin et al.

Understand

The current safeguard mechanisms for large language models (LLMs) are indeed susceptible to jailbreak attacks, making them inherently fragile.

  • Even the process of fine-tuning on apparently benign data for downstream tasks can jeopardize safety.
  • One potential solution is to conduct safety fine-tuning subsequent to downstream fine-tuning.
  • However, there's a risk of catastrophic forgetting during safety fine-tuning, where LLMs may regain safety measures but lose the task-specific knowledge acquired during downstream fine-tuning.

Reading the bibliography…