Fetching the paper…

EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model · Around