Fetching the paper…

How Visual Representations Map to Language Feature Space in Multimodal LLMs · Around