2018

Detecting Offensive Content in Open-domain Conversations using Two Stage Semi-supervision

Khatri, Chandra, Hedayatnia, Behnam, Goel, Rahul et al.

Understand

As open-ended human-chatbot interaction becomes commonplace, sensitive content detection gains importance.

  • In this work, we propose a two stage semi-supervised approach to bootstrap large-scale data for automatic sensitive language detection from publicly available web resources.
  • We explore various data selection methods including 1) using a blacklist to rank online discussion forums by the level of their sensitiveness followed by randomly sampling utterances and 2) training a weakly supervised model in conjunction with the blacklist for scoring sentences from online discussion forums to curate a dataset.
  • Our data collection strategy is flexible and allows the models to detect implicit sensitive content for which manual annotations may be difficult.

Reading the bibliography…