2020

Imitation Attacks and Defenses for Black-box Machine Translation Systems

Wallace, Eric, Stern, Mitchell, Song, Dawn

Understand

Adversaries may look to steal or attack black-box NLP systems, either for financial gain or to exploit model errors.

  • One setting of particular interest is machine translation (MT), where models have high commercial value and errors can be costly.
  • We investigate possible exploits of black-box MT systems and explore a preliminary defense against such threats.
  • We first show that MT systems can be stolen by querying them with monolingual sentences and training models to imitate their outputs.

Reading the bibliography…