2021

Just Say No: Analyzing the Stance of Neural Dialogue Generation in Offensive Contexts

Baheti, Ashutosh, Sap, Maarten, Ritter, Alan et al.

Understand

Dialogue models trained on human conversations inadvertently learn to generate toxic responses.

  • In addition to producing explicitly offensive utterances, these models can also implicitly insult a group or individual by aligning themselves with an offensive statement.
  • To better understand the dynamics of contextually offensive language, we investigate the stance of dialogue model responses in offensive Reddit conversations.
  • Specifically, we create ToxiChat, a crowd-annotated dataset of 2,000 Reddit threads and model responses labeled with offensive language and stance.

Reading the bibliography…