2020

DocVQA: A Dataset for VQA on Document Images

Mathew, Minesh, Karatzas, Dimosthenis, Jawahar, C. V.

Understand

We present a new dataset for Visual Question Answering (VQA) on document images called DocVQA.

  • The dataset consists of 50,000 questions defined on 12,000+ document images.
  • Detailed analysis of the dataset in comparison with similar datasets for VQA and reading comprehension is presented.
  • We report several baseline results by adopting existing VQA and reading comprehension models.

Reading the bibliography…