Fetching the paper…

Tencent Text-Video Retrieval: Hierarchical Cross-Modal Interactions with Multi-Level Representations · Around