Fetching the paper…

ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference · Around