Fetching the paper…

VideoBERT: A Joint Model for Video and Language Representation Learning · Around