Fetching the paper…

DOFA-CLIP: Multimodal Vision-Language Foundation Models for Earth Observation · Around