2016

A Semi-supervised Framework for Image Captioning

Chen, Wenhu, Lucchi, Aurelien, Hofmann, Thomas

Understand

State-of-the-art approaches for image captioning require supervised training data consisting of captions with paired image data.

  • These methods are typically unable to use unsupervised data such as textual data with no corresponding images, which is a much more abundant commodity.
  • We here propose a novel way of using such textual data by artificially generating missing visual information.
  • We evaluate this learning approach on a newly designed model that detects visual concepts present in an image and feed them to a reviewer-decoder architecture with an attention mechanism.

Reading the bibliography…