Fetching the paper…

Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training · Around