Fetching the paper…

MILES: Visual BERT Pre-training with Injected Language Semantics for Video-text Retrieval · Around