Fetching the paper…

Grounding Everything: Emerging Localization Properties in Vision-Language Transformers · Around