2023

Visual Adversarial Examples Jailbreak Aligned Large Language Models

Qi, Xiangyu, Huang, Kaixuan, Panda, Ashwinee et al.

Understand

Recently, there has been a surge of interest in integrating vision into Large Language Models (LLMs), exemplified by Visual Language Models (VLMs) such as Flamingo and GPT-4.

  • This paper sheds light on the security and safety implications of this trend.
  • First, we underscore that the continuous and high-dimensional nature of the visual input makes it a weak link against adversarial attacks, representing an expanded attack surface of vision-integrated LLMs.
  • Second, we highlight that the versatility of LLMs also presents visual attackers with a wider array of achievable adversarial objectives, extending the implications of security failures beyond mere misclassification.

Reading the bibliography…