Fetching the paper…
Reading the bibliography…
Large vision-language models (VLMs) have demonstrated remarkable capabilities in open-world multimodal understanding, yet their high computational overheads pose great challenges for practical deployment.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 1877–1901
1901
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , vol. 30, 2017
2017
Earlier work this paper cites.
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the v in vqa matter: Elevating the role of image understanding in visual question answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2017, pp. 6904–6913
2017
Earlier work this paper cites.
D. A. Hudson and C. D. Manning, “Gqa: A new dataset for real-world visual reasoning and compositional question answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019
2019
Earlier work this paper cites.
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach, “Towards vqa models that can read,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 8317–8326
2019
Earlier work this paper cites.
T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré, “FlashAttention: Fast and memory-efficient exact attention with IO-awareness,” in Advances in Neural Information Processing Systems (NeurIPS) , 2022
2022
Earlier work this paper cites.
P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, and A. Kalyan, “Learn to explain: Multimodal reasoning via thought chains for science question answering,” Advances in Neural Information Processing Systems , vol. 35, pp. 2507–2521, 2022
2022
Earlier work this paper cites.
OpenAI, “Gpt-4v(ision) system card,” OpenAI, Tech. Rep., 2023
2023
Earlier work this paper cites.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in Advances in Neural Information Processing Systems , vol. 36, 2023, pp. 34 892–34 916
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
H. Jiang, Q. Wu, C.-Y. Lin, Y. Yang, and L. Qiu, “Llmlingua: Compressing prompts for accelerated inference of large language models,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , 2023, pp. 13 358–13 376
2023
Earlier work this paper cites.
J. Mu, X. Li, and N. Goodman, “Learning to compress prompts with gist tokens,” Advances in Neural Information Processing Systems , vol. 36, pp. 19 327–19 352, 2023
2023
Earlier work this paper cites.
S. Dai, H. Genc, R. Venkatesan, and B. Khailany, “Efficient transformer inference with statically structured sparse attention,” in 2023 60th ACM/IEEE Design Automation Conference (DAC) . IEEE, 2023, pp. 1–6
2023
Earlier work this paper cites.
Z. Zhang, Y. Sheng, T. Zhou, T. Chen, L. Zheng, R. Cai, Z. Song, Y. Tian, C. Ré, C. Barrett et al. , “H2o: Heavy-hitter oracle for efficient generative inference of large language models,” Advances in Neural Information Processing Systems , vol. 36, pp. 34 661–34 710, 2023
2023
Earlier work this paper cites.
Y. Leviathan, M. Kalman, and Y. Matias, “Fast inference from transformers via speculative decoding,” in International Conference on Machine Learning . PMLR, 2023, pp. 19 274–19 286
2023
Earlier work this paper cites.
D. Bolya, C.-Y. Fu, X. Dai, P. Zhang, C. Feichtenhofer, and J. Hoffman, “Token merging: Your vit but faster,” in The Eleventh International Conference on Learning Representations , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
K. Mangalam, R. Akshulakov, and J. Malik, “Egoschema: A diagnostic benchmark for very long-form video language understanding,” in Advances in Neural Information Processing Systems , 2023
2023
Earlier work this paper cites.
Z. Chen, W. Wang, H. Tian, S. Ye, Z. Gao, E. Cui, W. Tong, K. Hu, J. Luo, Z. Ma et al. , “How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites,” Science China Information Sciences , vol. 67, no. 12, p. 220101, 2024
2024
Earlier work this paper cites.
Z. Peng, W. Wang, L. Dong, Y. Hao, S. Huang, S. Ma, Q. Ye, and F. Wei, “Grounding multimodal large language models to the world,” in The Twelfth International Conference on Learning Representations , 2024
2024
Earlier work this paper cites.
C. Ma, Y. Jiang, J. Wu, Z. Yuan, and X. Qi, “Groma: Localized visual tokenization for grounding multimodal large language models,” in European Conference on Computer Vision . Springer, 2024, pp. 417–435
2024
Earlier work this paper cites.
C. Luo, Y. Shen, Z. Zhu, Q. Zheng, Z. Yu, and C. Yao, “Layoutllm: Layout instruction tuning with large language models for document understanding,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2024, pp. 15 630–15 640
2024
Earlier work this paper cites.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
L. Chen, H. Zhao, T. Liu, S. Bai, J. Lin, C. Zhou, and B. Chang, “An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models,” in European Conference on Computer Vision . Springer, 2024, pp. 19–35
2024
Cited alongside, same era.
2024
2024
Later among the works it cites.
X. Miao, G. Oliaro, Z. Zhang, X. Cheng, Z. Wang, Z. Zhang, R. Y. Y. Wong, A. Zhu, L. Yang, X. Shi, C. Shi, Z. Chen, D. Arfeen, R. Cen, and Z. Jia, “Specinfer: Accelerating large language model serving with tree-based speculative inference and verification,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS) , 2024
2024
Later among the works it cites.
H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee, “Llava-next: Improved reasoning, ocr, and world knowledge,” January 2024. [Online]. Available: https://llava-vl.github.io/blog/2024-01-30-llava-next/
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
H. Liu, C. Li, Y. Li, and Y. J. Lee, “Improved baselines with visual instruction tuning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 26 296–26 306
2024
Cited alongside, same era.
A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,” in First Conference on Language Modeling , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
G. Xiao, Y. Tian, B. Chen, S. Han, and M. Lewis, “Efficient streaming language models with attention sinks,” in The Twelfth International Conference on Learning Representations , 2024
2024
Cited alongside, same era.
2024
Later among the works it cites.
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun et al. , “Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 9556–9567
2024
Later among the works it cites.
L. Chen, J. Li, X. Dong, P. Zhang, Y. Zang, Z. Chen, H. Duan, J. Wang, Y. Qiao, D. Lin, and F. Zhao, “Are we on the right way for evaluating large vision-language models?” in Advances in Neural Information Processing Systems , 2024
2024
Later among the works it cites.
Y. Liu, Z. Li, M. Huang, B. Yang, W. Yu, C. Li, X.-C. Yin, C.-L. Liu, L. Jin, and X. Bai, “Ocrbench: On the hidden mystery of ocr in large multimodal models,” Science China Information Sciences , vol. 67, no. 12, 2024
2024
Later among the works it cites.
X. Fu, Y. Hu, B. Li, Y. Feng, H. Wang, X. Lin, D. Roth, N. A. Smith, W.-C. Ma, and R. Krishna, “Blink: Multimodal large language models can see but not perceive,” in Proceedings of the European Conference on Computer Vision , 2024
2024
Later among the works it cites.
K. Li, Y. Wang, Y. He, Y. Li, Y. Wang, Y. Liu, Z. Wang, J. Xu, G. Chen, P. Lou, L. Wang, and Y. Qiao, “Mvbench: A comprehensive multi-modal video understanding benchmark,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 22 195–22 206
2024
Later among the works it cites.
2025
Closest in time.
Z. Shao, Z. Yu, J. Yu, X. Ouyang, L. Zheng, Z. Gai, M. Wang, Z. Kuang, and J. Ding, “Imp: Highly capable large multimodal models for mobile devices,” IEEE Transactions on Multimedia , 2025
2025
Closest in time.
Z. Shao, M. Wang, Z. Yu, W. Pan, Y. Yang, T. Wei, H. Zhang, N. Mao, W. Chen, and J. Yu, “Growing a twig to accelerate large vision-language models,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2025, pp. 20 064–20 074
2025
Closest in time.
Y. Sun, L. Dong, Y. Zhu, S. Huang, W. Wang, S. Ma, Q. Zhang, J. Wang, and F. Wei, “You only cache once: Decoder-decoder architectures for language models,” Advances in Neural Information Processing Systems , vol. 37, pp. 7339–7361, 2025
2025
Closest in time.
F. Liu, Y. Tang, Z. Liu, Y. Ni, D. Tang, K. Han, and Y. Wang, “Kangaroo: Lossless self-speculative decoding for accelerating llms via double early exiting,” Advances in Neural Information Processing Systems , vol. 37, pp. 11 946–11 965, 2025
2025
Closest in time.
Q. Zhang, A. Cheng, M. Lu, R. Zhang, Z. Zhuo, J. Cao, S. Guo, Q. She, and S. Zhang, “Beyond text-visual attention: Exploiting visual cues for effective token pruning in vlms,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2025, pp. 20 857–20 867
2025
Closest in time.
Y. Jiang, Q. Wu, W. Lin, W. Yu, and Y. Zhou, “What kind of visual tokens do we need? training-free visual token pruning for multi-modal large language models from the perspective of graph,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2025, aAAI 2025
2025
Closest in time.
M. Endo, X. Wang, and S. Yeung-Levy, “Feather the throttle: Revisiting visual token pruning for vision-language model acceleration,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2025, pp. 22 826–22 835
2025
Closest in time.
Z. Wen, Y. Gao, W. Li, C. He, and L. Zhang, “Token pruning in multimodal large language models: Are we solving the right problem?” in Findings of the Association for Computational Linguistics: ACL 2025 , 2025, pp. 15 537–15 549
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Y. Ji, J. Zhang, H. Xia, J. Chen, L. Shou, G. Chen, and H. Li, “SpecVLM: Enhancing speculative decoding of video LLMs via verifier-guided token pruning,” in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng, Eds. Suzhou, China: Association for Computational Linguistics, Nov. 2025, pp. 7205–7219. [Online]. Available: https://aclanthology.org/2025.emnlp-main.366/
2025
Closest in time.
J. Kang, H. Shu, W. Li, Y. Zhai, and X. Chen, “Vispec: Accelerating vision-language models with vision-aware speculative decoding,” in Advances in Neural Information Processing Systems 38 (NeurIPS 2025) , 2025. [Online]. Available: https://openreview.net/forum?id=x2BsIdJJJW
2025
Closest in time.
Y. Tang, S. Wang, L. Madaan, and R. Munos, “Beyond verifiable rewards: Scaling reinforcement learning in language models to unverifiable data,” in Advances in Neural Information Processing Systems , 2025. [Online]. Available: https://openreview.net/forum?id=pc6M9h3T9m
2025
Closest in time.
Y. Liu, H. Duan, Y. Zhang, B. Li, S. Zhang, W. Zhao, Y. Yuan, J. Wang, C. He, Z. Liu et al. , “Mmbench: Is your multi-modal model an all-around player?” in European Conference on Computer Vision . Springer, 2025, pp. 216–233
2025
Closest in time.
C. Fu, Y. Dai, Y. Luo, L. Li, S. Ren, R. Zhang, Z. Wang, C. Zhou, Y. Shen, M. Zhang et al. , “Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2025
2025
Closest in time.
X. Zhou, Z. Liu, A. Sims, H. Wang, T. Pang, C. Li, L. Wang, M. Lin, and C. Du, “Reinforcing general reasoning without verifiers,” in International Conference on Learning Representations , 2026. [Online]. Available: https://openreview.net/forum?id=nnwvwge40d
2026
Closest in time.