Fetching the paper…

Enhancing Vision Foundation Models via Multimodal Continual Pre-Training · Around