2020

Multimodal grid features and cell pointers for Scene Text Visual Question Answering

Gómez, Lluís, Biten, Ali Furkan, Tito, Rubèn et al.

Understand

This paper presents a new model for the task of scene text visual question answering, in which questions about a given image can only be answered by reading and understanding scene text that is present in it.

  • The proposed model is based on an attention mechanism that attends to multi-modal features conditioned to the question, allowing it to reason jointly about the textual and visual modalities in the scene.
  • The output weights of this attention module over the grid of multi-modal spatial features are interpreted as the probability that a certain spatial location of the image contains the answer text the to the given question.
  • Our experiments demonstrate competitive performance in two standard datasets.

Reading the bibliography…