Fetching the paper…
Reading the bibliography…
Recent advances in multimodal large language models (MLLMs) have expanded research in video understanding, primarily focusing on high-level tasks such as video captioning and question-answering.
Nothing clear enough to list yet.
Nothing clear enough to list yet.
Nothing clear enough to list yet.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…