Fetching the paper…

LAVENDER: Unifying Video-Language Understanding as Masked Language Modeling · Around