2015

Visual7W: Grounded Question Answering in Images

Zhu, Yuke, Groth, Oliver, Bernstein, Michael et al.

Understand

We have seen great progress in basic perceptual tasks such as object recognition and detection.

  • However, AI models still fail to match humans in high-level vision tasks due to the lack of capacities for deeper reasoning.
  • Recently the new task of visual question answering (QA) has been proposed to evaluate a model's capacity for deep image understanding.
  • Previous works have established a loose, global association between QA sentences and images.

Reading the bibliography…