2022

ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Hartvigsen, Thomas, Gabriel, Saadia, Palangi, Hamid et al.

Understand

Toxic language detection systems often falsely flag text that contains minority group mentions as toxic, as those groups are often the targets of online hate.

  • Such over-reliance on spurious correlations also causes systems to struggle with detecting implicitly toxic language.
  • To help mitigate these issues, we create ToxiGen, a new large-scale and machine-generated dataset of 274k toxic and benign statements about 13 minority groups.
  • We develop a demonstration-based prompting framework and an adversarial classifier-in-the-loop decoding method to generate subtly toxic and benign text with a massive pretrained language model.

Reading the bibliography…