2021

Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Wang, Boxin, Xu, Chejian, Wang, Shuohang et al.

Understand

Large-scale pre-trained language models have achieved tremendous success across a wide range of natural language understanding (NLU) tasks, even surpassing human performance.

  • However, recent studies reveal that the robustness of these models can be challenged by carefully crafted textual adversarial examples.
  • While several individual datasets have been proposed to evaluate model robustness, a principled and comprehensive benchmark is still missing.
  • In this paper, we present Adversarial GLUE (AdvGLUE), a new multi-task benchmark to quantitatively and thoroughly explore and evaluate the vulnerabilities of modern large-scale language models under various types of adversarial attacks.

Reading the bibliography…