Fetching the paper…

MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training · Around