Fetching the paper…

VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos · Around