2020

Reevaluating Adversarial Examples in Natural Language

Morris, John X., Lifland, Eli, Lanchantin, Jack et al.

Understand

State-of-the-art attacks on NLP models lack a shared definition of a what constitutes a successful attack.

  • We distill ideas from past work into a unified framework: a successful natural language adversarial example is a perturbation that fools the model and follows some linguistic constraints.
  • We then analyze the outputs of two state-of-the-art synonym substitution attacks.
  • We find that their perturbations often do not preserve semantics, and 38% introduce grammatical errors.

Reading the bibliography…