Fetching the paper…

LinVT: Empower Your Image-level Large Language Model to Understand Videos · Around