2023

Making Harmful Behaviors Unlearnable for Large Language Models

Zhou, Xin, Lu, Yi, Ma, Ruotian et al.

Understand

Large language models (LLMs) have shown great potential as general-purpose AI assistants in various domains.

  • To meet the requirements of different applications, LLMs are often customized by further fine-tuning.
  • However, the powerful learning ability of LLMs not only enables them to acquire new tasks but also makes them susceptible to learning undesired behaviors.
  • For example, even safety-aligned LLMs can be easily fine-tuned into harmful assistants as the fine-tuning data often contains implicit or explicit harmful content.

Reading the bibliography…