2018

TVQA: Localized, Compositional Video Question Answering

Lei, Jie, Yu, Licheng, Bansal, Mohit et al.

Understand

Recent years have witnessed an increasing interest in image-based question-answering (QA) tasks.

  • However, due to data limitations, there has been much less work on video-based QA.
  • In this paper, we present TVQA, a large-scale video QA dataset based on 6 popular TV shows.
  • TVQA consists of 152,545 QA pairs from 21,793 clips, spanning over 460 hours of video.

Reading the bibliography…