2017

MUTAN: Multimodal Tucker Fusion for Visual Question Answering

Ben-younes, Hedi, Cadene, Rémi, Cord, Matthieu et al.

Understand

Bilinear models provide an appealing framework for mixing and merging information in Visual Question Answering (VQA) tasks.

  • They help to learn high level associations between question meaning and visual concepts in the image, but they suffer from huge dimensionality issues.
  • We introduce MUTAN, a multimodal tensor-based Tucker decomposition to efficiently parametrize bilinear interactions between visual and textual representations.
  • Additionally to the Tucker framework, we design a low-rank matrix-based decomposition to explicitly constrain the interaction rank.

Reading the bibliography…