Fetching the paper…

SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding · Around