2022

Evaluating Psychological Safety of Large Language Models

Li, Xingxuan, Li, Yutong, Qiu, Lin et al.

Understand

In this work, we designed unbiased prompts to systematically evaluate the psychological safety of large language models (LLMs).

  • First, we tested five different LLMs by using two personality tests: Short Dark Triad (SD-3) and Big Five Inventory (BFI).
  • All models scored higher than the human average on SD-3, suggesting a relatively darker personality pattern.
  • Despite being instruction fine-tuned with safety metrics to reduce toxicity, InstructGPT, GPT-3.5, and GPT-4 still showed dark personality patterns; these models scored higher than self-supervised GPT-3 on the Machiavellianism and narcissism traits on SD-3.

Reading the bibliography…