Fetching the paper…
Reading the bibliography…
Video Large Language Models (VLLMs) excel in video understanding, but their excessive visual tokens pose a significant computational challenge for real-world applications.
A simple and effective algorithm for the MaxMin diversity problem
Porumbel, D. C.; Hao, J.-K.; and Glover, F. 2011 · 2011
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Measuring diversity. A review and an empirical analysis
Parreño, F.; Álvarez-Valdés, R.; and Martí, R. 2021 · 2021
Earlier work this paper cites.
Filter, correlate, compress: Training-free token reduction for mllm acceleration
Han, Y.; Liu, X.; Zhang, Z.; Ding, P.; Wang, D.; Chen, H.; Yan, Q.; and Huang, S. 2024 · 2024
Earlier work this paper cites.
Prunevid: Visual token pruning for efficient video large language models
Huang, X.; Zhou, H.; and Han, K. 2024 · 2024
Earlier work this paper cites.
Llama-vid: An image is worth 2 tokens in large language models
Li, Y.; Wang, C.; and Jia, J. 2024 · 2024
Earlier work this paper cites.
Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Shang, Y.; Cai, M.; Xu, B.; Lee, Y. J.; and Yan, Y. 2024 · 2024
Cited alongside, same era.
Moviechat: From dense token to sparse memory for long video understanding
Song, E.; Chai, W.; Wang, G.; Zhang, Y.; Zhou, H.; Wu, F.; Chi, H.; Guo, X.; Ye, T.; Zhang, Y.; et al. 2024 · 2024
Cited alongside, same era.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Wang, P.; Bai, S.; Tan, S.; Wang, S.; Fan, Z.; Bai, J.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; et al. 2024 · 2024
Cited alongside, same era.
Longvideobench: A benchmark for long-context interleaved video-language understanding
Wu, H.; Li, D.; Chen, B.; and Li, J. 2024 · 2024
Cited alongside, same era.
Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy reduction
Divprune: Diversity-based visual token pruning for large multimodal models
Alvar, S. R.; Singh, G.; Akbari, M.; and Zhang, Y. 2025 · 2025
Closest in time.
HoliTom: Holistic Token Merging for Fast Video Large Language Models
Shao, K.; Tao, K.; Qin, C.; You, H.; Sui, Y.; and Wang, H. 2025 · 2025
Closest in time.
Fastvid: Dynamic density pruning for fast video large language models
Shen, L.; Gong, G.; He, T.; Zhang, Y.; Liu, P.; Zhao, S.; and Ding, G. 2025 · 2025
Closest in time.
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
Sun, B.; Zhao, J.; Wei, X.; and Hou, Q. 2025 · 2025
Closest in time.
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
Tao, K.; Qin, C.; You, H.; Sui, Y.; and Wang, H. 2025 · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xing, L.; Huang, Q.; Dong, X.; Lu, J.; Zhang, P.; Zang, Y.; Cao, Y.; He, C.; Wang, J.; Wu, F.; et al. 2024 · 2024
Cited alongside, same era.
Mlvu: A comprehensive benchmark for multi-task long video understanding
Zhou, J.; Shu, Y.; Zhao, B.; Wu, B.; Xiao, S.; Yang, X.; Xiong, Y.; Zhang, B.; Huang, T.; and Liu, Z. 2024 · 2024
Cited alongside, same era.
An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Chen, L.; Zhao, H.; Liu, T.; Bai, S.; Lin, J.; Zhou, C.; and Chang, B. 2024a
Cited in the paper.
Longvila: Scaling long-context visual language models for long videos
Chen, Y.; Xue, F.; Li, D.; Hu, Q.; Zhu, L.; Li, X.; Fang, Y.; Tang, H.; Yang, S.; Liu, Z.; et al. 2024b
Cited in the paper.
Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Fu, C.; Dai, Y.; Luo, Y.; Li, L.; Ren, S.; Zhang, R.; Wang, Z.; Zhou, C.; Shen, Y.; Zhang, M.; et al. 2024a
Cited in the paper.
Fu, T.; Liu, T.; Han, Q.; Dai, G.; Yan, S.; Yang, H.; Ning, X.; and Wang, Y. 2024b
Cited in the paper.
Llava-onevision: Easy visual task transfer
Li, B.; Zhang, Y.; Guo, D.; Zhang, R.; Li, F.; Zhang, H.; Zhang, K.; Zhang, P.; Li, Y.; Liu, Z.; et al. 2024a
Cited in the paper.
Mvbench: A comprehensive multi-modal video understanding benchmark
Li, K.; Wang, Y.; He, Y.; Li, Y.; Wang, Y.; Liu, Y.; Wang, Z.; Xu, J.; Chen, G.; Luo, P.; et al. 2024b
Cited in the paper.
Closest in time.
Voco-llama: Towards vision compression with large language models
Ye, X.; Gan, Y.; Huang, X.; Ge, Y.; and Tang, Y. 2025 · 2025
Closest in time.