2020

Entangled Watermarks as a Defense against Model Extraction

Jia, Hengrui, Choquette-Choo, Christopher A., Chandrasekaran, Varun et al.

Understand

Machine learning involves expensive data collection and training procedures.

  • Model owners may be concerned that valuable intellectual property can be leaked if adversaries mount model extraction attacks.
  • As it is difficult to defend against model extraction without sacrificing significant prediction accuracy, watermarking instead leverages unused model capacity to have the model overfit to outlier input-output pairs.
  • Such pairs are watermarks, which are not sampled from the task distribution and are only known to the defender.

Reading the bibliography…