Fetching the paper…

COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning · Around