2024

GUARD: Role-playing to Generate Natural-language Jailbreakings to Test Guideline Adherence of Large Language Models

Jin, Haibo, Chen, Ruoxi, Zhang, Peiyan et al.

Understand

The discovery of "jailbreaks" to bypass safety filters of Large Language Models (LLMs) and harmful responses have encouraged the community to implement safety measures.

  • One major safety measure is to proactively test the LLMs with jailbreaks prior to the release.
  • Therefore, such testing will require a method that can generate jailbreaks massively and efficiently.
  • In this paper, we follow a novel yet intuitive strategy to generate jailbreaks in the style of the human generation.

Reading the bibliography…