Fetching the paper…
Reading the bibliography…
Multi-modal Large Language Models (MLLMs) struggle with long videos due to the need for excessive visual tokens.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Next-qa: Next phase of question-answering to explaining temporal actions
Xiao, J., Shang, X., Yao, A., and Chua, T.-S · 2021
Earlier work this paper cites.
Himakunthala, V., Ouyang, A., Rose, D., He, R., Mei, A., Lu, Y., Sonar, C., Saxon, M., and Wang, W. Y · 2023
Earlier work this paper cites.
Video-llava: Learning united visual representation by alignment before projection
Lin, B., Ye, Y., Zhu, B., Cui, J., Ning, M., Jin, P., and Yuan, L · 2023
Earlier work this paper cites.
Video-chatgpt: Towards detailed video understanding via large vision and language models
Maaz, M., Rasheed, H., Khan, S., and Khan, F. S · 2023
Earlier work this paper cites.
Gpt-4 technical report
OpenAI · 2023
Earlier work this paper cites.
Chatgpt: Optimizing language models for dialogue
OpenAI · 2023
Earlier work this paper cites.
Chatvideo: A tracklet-centric multimodal and versatile video understanding system
Wang, J., Chen, D., Luo, C., Dai, X., Yuan, L., Wu, Z., and Jiang, Y.-G · 2023
Earlier work this paper cites.
mplug-owl: Modularization empowers large language models with multimodality
Ye, Q., Xu, H., Xu, G., Ye, J., Yan, M., Zhou, Y., Wang, J., Hu, A., Shi, P., Shi, Y., et al · 2023
Earlier work this paper cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2023
Earlier work this paper cites.
https://www.anthropic.com/news/claude-3-family
Anthropic · 2024
Earlier work this paper cites.
Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms
Cheng, Z., Leng, S., Zhang, H., Xin, Y., Li, X., Chen, G., Zhu, Y., Zhang, W., Luo, Z., Zhao, D., et al · 2024
Cited alongside, same era.
Han, S., Huang, W., Shi, H., Zhuo, L., Su, X., Zhang, S., Zhou, X., Qi, X., Liao, Y., and Liu, S · 2024
Cited alongside, same era.
Ma-lmm: Memory-augmented large multimodal model for long-term video understanding
He, B., Li, H., Jang, Y. K., Jia, M., Cao, X., Shah, A., Shrivastava, A., and Lim, S.-N · 2024
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Cited alongside, same era.
Moviechat: From dense token to sparse memory for long video understanding
Song, E., Chai, W., Wang, G., Zhang, Y., Zhou, H., Wu, F., Chi, H., Guo, X., Ye, T., Zhang, Y., et al · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Team, G., Georgiev, P., Lei, V. I., Burnell, R., Bai, L., Gulati, A., Tanzer, G., Vincent, D., Pan, Z., Wang, S., et al · 2024
Later among the works it cites.
Longvideobench: A benchmark for long-context interleaved video-language understanding
Wu, H., Li, D., Chen, B., and Li, J · 2024
Later among the works it cites.
Pllava: Parameter-free llava extension from images to videos for video dense captioning
Xu, L., Zhao, Y., Zhou, D., Lin, Z., Ng, S. K., and Feng, J · 2024
Later among the works it cites.
Longvila: Scaling long-context visual language models for long videos
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jin, P., Takanobu, R., Zhang, W., Cao, X., and Yuan, L · 2024
Cited alongside, same era.
An image grid can be worth a video: Zero-shot video question answering using a vlm
Kim, W., Choi, C., Lee, W., and Rhee, W · 2024
Cited alongside, same era.
Mvbench: A comprehensive multi-modal video understanding benchmark
Li, K., Wang, Y., He, Y., Li, Y., Wang, Y., Liu, Y., Wang, Z., Xu, J., Chen, G., Luo, P., et al · 2024
Cited alongside, same era.
Openai, gpt-40
OpenAI · 2024
Cited alongside, same era.
Timechat: A time-sensitive multimodal large language model for long video understanding
Ren, S., Yao, L., Li, S., Sun, X., and Hou, L · 2024
Cited alongside, same era.
Unlocking video-llm via agent-of-thoughts distillation
Shi, Y., Di, S., Chen, Q., and Xie, W · 2024
Cited alongside, same era.
Video-xl: Extra-long vision language model for hour-scale video understanding
Shu, Y., Zhang, P., Liu, Z., Qin, M., Zhou, J., Huang, T., and Zhao, B · 2024
Cited alongside, same era.
Sharegpt4video: Improving video understanding and generation with better captions
Chen, L., Wei, X., Li, J., Dong, X., Zhang, P., Zang, Y., Chen, Z., Duan, H., Lin, B., Tang, Z., et al
Cited in the paper.
Xue, F., Chen, Y., Li, D., Hu, Q., Zhu, L., Li, X., Fang, Y., Tang, H., Yang, S., Liu, Z., et al · 2024
Later among the works it cites.
Llava-next: A strong zero-shot video understanding model, April 2024b
Zhang, Y., Li, B., Liu, h., Lee, Y. j., Gui, L., Fu, D., Feng, J., Liu, Z., and Li, C · 2024
Later among the works it cites.
Videogen-of-thought: A collaborative framework for multi-shot video generation
Zheng, M., Xu, Y., Huang, H., Ma, X., Liu, Y., Shu, W., Pang, Y., Tang, F., Chen, Q., Yang, H., et al · 2024
Later among the works it cites.
Mlvu: A comprehensive benchmark for multi-task long video understanding
Zhou, J., Shu, Y., Zhao, B., Wu, B., Xiao, S., Yang, X., Xiong, Y., Zhang, B., Huang, T., and Liu, Z · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Closest in time.
Llama-vid: An image is worth 2 tokens in large language models
Li, Y., Wang, C., and Jia, J · 2025
Closest in time.
St-llm: Large language models are effective temporal learners
Liu, R., Li, C., Tang, H., Ge, Y., Shan, Y., and Li, G · 2025
Closest in time.