Fetching the paper…

Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models · Around