Fetching the paper…
Reading the bibliography…
Recently, multimodal large language models (MM-LLMs) have achieved significant success in various tasks, but their high computational costs limit widespread application.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
Towards vqa models that can read
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Dynamicvit: Efficient vision transformers with dynamic token sparsification
Rao, Y., Zhao, W., Liu, B., Lu, J., Zhou, J., and Hsieh, C.-J · 2021
Earlier work this paper cites.
Spvit: Enabling faster vision transformers via latency-aware soft token pruning
Kong, Z., Dong, P., Ma, X., Meng, X., Niu, W., Sun, M., Shen, X., Yuan, G., Ren, B., Tang, H., et al · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Lu, P., Mishra, S., Xia, T., Qiu, L., Chang, K.-W., Zhu, S.-C., Tafjord, O., Clark, P., and Kalyan, A · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Token merging: Your vit but faster
Bolya, D., Fu, C.-Y., Dai, X., Zhang, P., Feichtenhofer, C., and Hoffman, J · 2023
Earlier work this paper cites.
Diffrate: Differentiable compression rate for efficient vision transformers
Chen, M., Shao, W., Xu, P., Lin, M., Zhang, K., Chao, F., Ji, R., Qiao, Y., and Luo, P · 2023
Earlier work this paper cites.
Mobilevlm: A fast, reproducible and strong vision language assistant for mobile devices
Chu, X., Qiao, L., Lin, X., Xu, S., Yang, Y., Hu, Y., Wei, F., Zhang, X., Zhang, B., Wei, X., et al · 2023
Earlier work this paper cites.
Mme: A comprehensive evaluation benchmark for multimodal large language models
Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Yang, J., Zheng, X., Li, K., Sun, X., Wu, Y., and Ji, R · 2023
Cited alongside, same era.
Evaluating object hallucination in large vision-language models
Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, W. X., and Wen, J.-R · 2023
Cited alongside, same era.
Improved baselines with visual instruction tuning
Liu, H., Li, C., Li, Y., and Lee, Y. J · 2023
Cited alongside, same era.
Beyond attentive tokens: Incorporating token importance and diversity for efficient vision transformers
Long, S., Zhao, Z., Pi, J., Wang, S., and Wang, J · 2023
Cited alongside, same era.
Gpt-4v (ision) system card
OpenAI · 2023
Cited alongside, same era.
Rethinking token reduction in mllms: Towards a unified paradigm for training-free acceleration
Han, Y., Liu, X., Ding, P., Wang, D., Chen, H., Yan, Q., and Huang, S · 2024
Closest in time.
Multi-criteria token fusion with one-step-ahead attention for efficient vision transformers
Lee, S., Choi, J., and Kim, H. J · 2024
Closest in time.
Snapkv: Llm knows what you are looking for before generation
Li, Y., Huang, Y., Yang, B., Venkitesh, B., Locatelli, A., Ye, H., Cai, T., Lewis, P., and Chen, D · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024a
Liu, H., Li, C., Li, Y., Li, B., Zhang, Y., Shen, S., and Lee, Y. J · 2024
Closest in time.
Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Shang, Y., Cai, M., Xu, B., Lee, Y. J., and Yan, Y · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Cited alongside, same era.
Ppt: Token pruning and pooling for efficient vision transformers
Wu, X., Zeng, F., Wang, X., and Chen, X · 2023
Cited alongside, same era.
No token left behind: Efficient vision transformer via dynamic token idling
Xu, X., Li, C., Chen, Y., Chang, X., Liu, J., and Wang, S · 2023
Cited alongside, same era.
Mm-vet: Evaluating large multimodal models for integrated capabilities
Yu, W., Yang, Z., Li, L., Wang, J., Lin, K., Liu, Z., Wang, X., and Wang, L · 2023
Cited alongside, same era.
Tinygpt-v: Efficient multimodal large language model via small backbones
Yuan, Z., Li, Z., and Sun, L · 2023
Cited alongside, same era.
Feather the throttle: Revisiting visual token pruning for vision-language model acceleration
Endo, M., Wang, X., and Yeung-Levy, S · 2024
Cited alongside, same era.
Closest in time.
Crossget: Cross-guided ensemble of tokens for accelerating vision-language transformers
Shi, D., Tao, C., Rao, A., Yang, Z., Yuan, C., and Wang, J · 2024
Closest in time.
SmartTrim: Adaptive tokens and attention pruning for efficient vision-language models
Wang, Z., Chen, J., Zhou, W., Zhu, H., Liang, J., Shan, L., Liu, M., Xu, D., Yang, Q., and Qin, B · 2024
Closest in time.
Llm inference unveiled: Survey and roofline model insights
Yuan, Z., Shang, Y., Zhou, Y., Dong, Z., Xue, C., Wu, B., Li, Z., Gu, Q., Lee, Y. J., Yan, Y., et al · 2024
Closest in time.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Ré, C., Barrett, C., et al · 2024
Closest in time.
A survey on efficient inference for large language models
Zhou, Z., Ning, X., Hong, K., Fu, T., Xu, J., Li, S., Lou, Y., Wang, L., Yuan, Z., Li, X., et al · 2024
Closest in time.
Recoverable compression: A multimodal vision token recovery mechanism guided by text information
Chen, Y., Xu, J., Zhang, X.-Y., Liu, W.-Z., Liu, Y.-Y., and Liu, C.-L · 2025
Closest in time.