Fetching the paper…

ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification · Around