2016

Hierarchical Question-Image Co-Attention for Visual Question Answering

Lu, Jiasen, Yang, Jianwei, Batra, Dhruv et al.

Understand

A number of recent works have proposed attention models for Visual Question Answering (VQA) that generate spatial maps highlighting image regions relevant to answering the question.

  • In this paper, we argue that in addition to modeling "where to look" or visual attention, it is equally important to model "what words to listen to" or question attention.
  • We present a novel co-attention model for VQA that jointly reasons about image and question attention.
  • In addition, our model reasons about the question (and consequently the image via the co-attention mechanism) in a hierarchical fashion via a novel 1-dimensional convolution neural networks (CNN).

Reading the bibliography…