2016

Image Captioning and Visual Question Answering Based on Attributes and External Knowledge

Wu, Qi, Shen, Chunhua, Hengel, Anton van den et al.

Understand

Much recent progress in Vision-to-Language problems has been achieved through a combination of Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs).

  • This approach does not explicitly represent high-level semantic concepts, but rather seeks to progress directly from image features to text.
  • In this paper we first propose a method of incorporating high-level concepts into the successful CNN-RNN approach, and show that it achieves a significant improvement on the state-of-the-art in both image captioning and visual question answering.
  • We further show that the same mechanism can be used to incorporate external knowledge, which is critically important for answering high level visual questions.

Reading the bibliography…