2019

RUBi: Reducing Unimodal Biases in Visual Question Answering

Cadene, Remi, Dancette, Corentin, Ben-younes, Hedi et al.

Understand

Visual Question Answering (VQA) is the task of answering questions about an image.

  • Some VQA models often exploit unimodal biases to provide the correct answer without using the image information.
  • As a result, they suffer from a huge drop in performance when evaluated on data outside their training set distribution.
  • This critical issue makes them unsuitable for real-world settings.

Reading the bibliography…