2016

Dual Attention Networks for Multimodal Reasoning and Matching

Nam, Hyeonseob, Ha, Jung-Woo, Kim, Jeonghee

Understand

We propose Dual Attention Networks (DANs) which jointly leverage visual and textual attention mechanisms to capture fine-grained interplay between vision and language.

  • DANs attend to specific regions in images and words in text through multiple steps and gather essential information from both modalities.
  • Based on this framework, we introduce two types of DANs for multimodal reasoning and matching, respectively.
  • The reasoning model allows visual and textual attentions to steer each other during collaborative inference, which is useful for tasks such as Visual Question Answering (VQA).

Reading the bibliography…