2021

T2VLAD: Global-Local Sequence Alignment for Text-Video Retrieval

Wang, Xiaohan, Zhu, Linchao, Yang, Yi

Understand

Text-video retrieval is a challenging task that aims to search relevant video contents based on natural language descriptions.

  • The key to this problem is to measure text-video similarities in a joint embedding space.
  • However, most existing methods only consider the global cross-modal similarity and overlook the local details.
  • Some works incorporate the local comparisons through cross-modal local matching and reasoning.

Reading the bibliography…