Fetching the paper…

VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending · Around