Fetching the paper…

Unified Multimodal Pre-training and Prompt-based Tuning for Vision-Language Understanding and Generation · Around