2023

Removing RLHF Protections in GPT-4 via Fine-Tuning

Zhan, Qiusi, Fang, Richard, Bindu, Rohan et al.

Understand

As large language models (LLMs) have increased in their capabilities, so does their potential for dual use.

  • To reduce harmful outputs, produces and vendors of LLMs have used reinforcement learning with human feedback (RLHF).
  • In tandem, LLM vendors have been increasingly enabling fine-tuning of their most powerful models.
  • However, concurrent work has shown that fine-tuning can remove RLHF protections.

Reading the bibliography…