2015

Jointly Modeling Embedding and Translation to Bridge Video and Language

Pan, Yingwei, Mei, Tao, Yao, Ting et al.

Understand

Automatically describing video content with natural language is a fundamental challenge of multimedia.

  • Recurrent Neural Networks (RNN), which models sequence dynamics, has attracted increasing attention on visual interpretation.
  • However, most existing approaches generate a word locally with given previous words and the visual content, while the relationship between sentence semantics and visual content is not holistically exploited.
  • As a result, the generated sentences may be contextually correct but the semantics (e.g., subjects, verbs or objects) are not true.

Reading the bibliography…