2023

Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements

Deng, Jiawen, Cheng, Jiale, Sun, Hao et al.

Understand

As generative large model capabilities advance, safety concerns become more pronounced in their outputs.

  • To ensure the sustainable growth of the AI ecosystem, it's imperative to undertake a holistic evaluation and refinement of associated safety risks.
  • This survey presents a framework for safety research pertaining to large models, delineating the landscape of safety risks as well as safety evaluation and improvement methods.
  • We begin by introducing safety issues of wide concern, then delve into safety evaluation methods for large models, encompassing preference-based testing, adversarial attack approaches, issues detection, and other advanced evaluation methods.

Reading the bibliography…