Fetching the paper…
Reading the bibliography…
Contrastive learning-based video-language representation learning approaches, e.g., CLIP, have achieved outstanding performance, which pursue semantic interaction upon pre-defined video-text pairs.
Nothing clear enough to list yet.
Nothing clear enough to list yet.
Nothing clear enough to list yet.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…