2018

A Visual Attention Grounding Neural Model for Multimodal Machine Translation

Zhou, Mingyang, Cheng, Runxiang, Lee, Yong Jae et al.

Understand

We introduce a novel multimodal machine translation model that utilizes parallel visual and textual information.

  • Our model jointly optimizes the learning of a shared visual-language embedding and a translator.
  • The model leverages a visual attention grounding mechanism that links the visual semantics with the corresponding textual semantics.
  • Our approach achieves competitive state-of-the-art results on the Multi30K and the Ambiguous COCO datasets.

Reading the bibliography…