Fetching the paper…

Masked Vision and Language Modeling for Multi-modal Representation Learning · Around