Fetching the paper…

Towards a Unified Foundation Model: Jointly Pre-Training Transformers on Unpaired Images and Text · Around