2016

Dense Captioning with Joint Inference and Visual Context

Yang, Linjie, Tang, Kevin, Yang, Jianchao et al.

Understand

Dense captioning is a newly emerging computer vision topic for understanding images with dense language descriptions.

  • The goal is to densely detect visual concepts (e.g., objects, object parts, and interactions between them) from images, labeling each with a short descriptive phrase.
  • We identify two key challenges of dense captioning that need to be properly addressed when tackling the problem.
  • First, dense visual concept annotations in each image are associated with highly overlapping target regions, making accurate localization of each visual concept challenging.

Reading the bibliography…