2020

Universal Adversarial Attacks with Natural Triggers for Text Classification

Song, Liwei, Yu, Xinwei, Peng, Hsuan-Tung et al.

Understand

Recent work has demonstrated the vulnerability of modern text classifiers to universal adversarial attacks, which are input-agnostic sequences of words added to text processed by classifiers.

  • Despite being successful, the word sequences produced in such attacks are often ungrammatical and can be easily distinguished from natural text.
  • We develop adversarial attacks that appear closer to natural English phrases and yet confuse classification systems when added to benign inputs.
  • We leverage an adversarially regularized autoencoder (ARAE) to generate triggers and propose a gradient-based search that aims to maximize the downstream classifier's prediction loss.

Reading the bibliography…