Fetching the paper…

Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens · Around