2020

On Modality Bias in the TVQA Dataset

Winterbottom, Thomas, Xiao, Sarah, McLean, Alistair et al.

Understand

TVQA is a large scale video question answering (video-QA) dataset based on popular TV shows.

  • The questions were specifically designed to require "both vision and language understanding to answer".
  • In this work, we demonstrate an inherent bias in the dataset towards the textual subtitle modality.
  • We infer said bias both directly and indirectly, notably finding that models trained with subtitles learn, on-average, to suppress video feature contribution.

Reading the bibliography…