Fetching the paper…

Introducing Visual Perception Token into Multimodal Large Language Model · Around