2022

CommonsenseQA 2.0: Exposing the Limits of AI through Gamification

Talmor, Alon, Yoran, Ori, Bras, Ronan Le et al.

Understand

Constructing benchmarks that test the abilities of modern natural language understanding models is difficult - pre-trained language models exploit artifacts in benchmarks to achieve human parity, but still fail on adversarial examples and make errors that demonstrate a lack of common sense.

  • In this work, we propose gamification as a framework for data construction.
  • The goal of players in the game is to compose questions that mislead a rival AI while using specific phrases for extra points.
  • The game environment leads to enhanced user engagement and simultaneously gives the game designer control over the collected data, allowing us to collect high-quality data at scale.

Reading the bibliography…