2023

DEPN: Detecting and Editing Privacy Neurons in Pretrained Language Models

Wu, Xinwei, Li, Junzhuo, Xu, Minghui et al.

Understand

Large language models pretrained on a huge amount of data capture rich knowledge and information in the training data.

  • The ability of data memorization and regurgitation in pretrained language models, revealed in previous studies, brings the risk of data leakage.
  • In order to effectively reduce these risks, we propose a framework DEPN to Detect and Edit Privacy Neurons in pretrained language models, partially inspired by knowledge neurons and model editing.
  • In DEPN, we introduce a novel method, termed as privacy neuron detector, to locate neurons associated with private information, and then edit these detected privacy neurons by setting their activations to zero.

Reading the bibliography…