Fetching the paper…
Reading the bibliography…
Multimodal large language models (MLLMs) demand considerable computations for inference due to the extensive parameters and the additional input tokens needed for visual information representation.
A Diagram is Worth A Dozen Images
Kembhavi, A.; Salvato, M.; Kolve, E.; Seo, M.; Hajishirzi, H.; and Farhadi, A. 2016 · 2016
Earlier work this paper cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Jang, Y.; Song, Y.; Yu, Y.; Kim, Y.; and Kim, G. 2017 · 2017
Earlier work this paper cites.
Token Pooling in Vision Transformers
Marin, D.; Chang, J.-H. R.; Ranjan, A.; Prabhu, A.; Rastegari, M.; and Tuzel, O. 2021 · 2021
Earlier work this paper cites.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Earlier work this paper cites.
Dynamicvit: Efficient Vision Transformers with Dynamic Token Sparsification
Rao, Y.; Zhao, W.; Liu, B.; Lu, J.; Zhou, J.; and Hsieh, C.-J. 2021 · 2021
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T.; Fu, D.; Ermon, S.; Rudra, A.; and Ré, C. 2022 · 2022
Earlier work this paper cites.
EViT: Expediting Vision Transformers via Token Reorganizations
Liang, Y.; Chongjian, G.; Tong, Z.; Song, Y.; Wang, J.; and Xie, P. 2022 · 2022
Earlier work this paper cites.
Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Lu, P.; Mishra, S.; Xia, T.; Qiu, L.; Chang, K.-W.; Zhu, S.-C.; Tafjord, O.; Clark, P.; and Kalyan, A. 2022 · 2022
Earlier work this paper cites.
Evo-Vit: Slow-Fast Token Evolution for Dynamic Vision Transformer
Xu, Y.; Zhang, Z.; Zhang, M.; Sheng, K.; Li, K.; Dong, W.; Zhang, L.; Xu, C.; and Sun, X. 2022 · 2022
Earlier work this paper cites.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Earlier work this paper cites.
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Bai, J.; Bai, S.; Yang, S.; Wang, S.; Tan, S.; Wang, P.; Lin, J.; Zhou, C.; and Zhou, J. 2023 · 2023
Cited alongside, same era.
Token Merging: Your ViT But Faster
Bolya, D.; Fu, C.-Y.; Dai, X.; Zhang, P.; Feichtenhofer, C.; and Hoffman, J. 2023 · 2023
Cited alongside, same era.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; Stoica, I.; and Xing, E. P. 2023 · 2023
Cited alongside, same era.
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Fu, C.; Chen, P.; Shen, Y.; Qin, Y.; Zhang, M.; Lin, X.; Yang, J.; Zheng, X.; Li, K.; Sun, X.; Wu, Y.; and Ji, R. 2023 · 2023
Cited alongside, same era.
Efficient Streaming Language Models with Attention Sinks
Xiao, G.; Tian, Y.; Chen, B.; Han, S.; and Lewis, M. 2023 · 2023
Later among the works it cites.
Mmmu: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert Agi
Yue, X.; Ni, Y.; Zhang, K.; Zheng, T.; Liu, R.; Zhang, G.; Stevens, S.; Jiang, D.; Ren, W.; Sun, Y.; et al. 2023 · 2023
Later among the works it cites.
LMMs-Eval: Accelerating the Development of Large Multimoal Models
Bo Li, K. Z., Peiyuan Zhang; et al. 2024 · 2024
Closest in time.
Cao, J.; Ye, P.; Li, S.; Yu, C.; Tang, Y.; Lu, J.; and Chen, T. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lai, X.; Tian, Z.; Chen, Y.; Li, Y.; Yuan, Y.; Liu, S.; and Jia, J. 2023 · 2023
Cited alongside, same era.
LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Li, Y.; Wang, C.; and Jia, J. 2023 · 2023
Cited alongside, same era.
Efficiently Scaling Transformer Inference
Pope, R.; Douglas, S.; Chowdhery, A.; Devlin, J.; Bradbury, J.; Heek, J.; Xiao, K.; Agrawal, S.; and Dean, J. 2023 · 2023
Cited alongside, same era.
Crossget: Cross-guided Ensemble of Tokens for Accelerating Vision-Language Transformers
Shi, D.; Tao, C.; Rao, A.; Yang, Z.; Yuan, C.; and Wang, J. 2023 · 2023
Cited alongside, same era.
Gemini: A Family Of Highly Capable Multimodal Models
Team, G.; Anil, R.; Borgeaud, S.; Wu, Y.; Alayrac, J.-B.; Yu, J.; Soricut, R.; Schalkwyk, J.; Dai, A. M.; Hauth, A.; et al. 2023 · 2023
Cited alongside, same era.
PPT: Token Pruning and Pooling for Efficient Vision Transformers
Wu, X.; Zeng, F.; Wang, X.; Wang, Y.; and Chen, X. 2023 · 2023
Cited alongside, same era.
CF-ViT: A General Coarse-to-Fine Method for Vision Transformer
Chen, M.; Lin, M.; Li, K.; Shen, Y.; Wu, Y.; Chao, F.; and Ji, R. 2023a
Cited in the paper.
DiffRate : Differentiable Compression Rate for Efficient Vision Transformers
Chen, M.; Shao, W.; Xu, P.; Lin, M.; Zhang, K.; Chao, F.; Ji, R.; Qiao, Y.; and Luo, P. 2023b
Cited in the paper.
Chen, L.; Zhao, H.; Liu, T.; Bai, S.; Lin, J.; Zhou, C.; and Chang, B. 2024 · 2024
Closest in time.
VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models
Duan, H.; Yang, J.; Qiao, Y.; Fang, X.; Chen, L.; Liu, Y.; Dong, X.; Zang, Y.; Zhang, P.; Wang, J.; Lin, D.; and Chen, K. 2024 · 2024
Closest in time.
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Fu, C.; Dai, Y.; Luo, Y.; Li, L.; Ren, S.; Zhang, R.; Wang, Z.; Zhou, C.; Shen, Y.; Zhang, M.; et al. 2024 · 2024
Closest in time.
Llava-uhd: an lmm perceiving any aspect ratio and high-resolution images
Guo, Z.; Xu, R.; Yao, Y.; Cui, J.; Ni, Z.; Ge, C.; Chua, T.-S.; Liu, Z.; and Huang, G. 2024 · 2024
Closest in time.
LLaVA-NeXT: Improved Reasoning, OCR, and World Knowledge
Liu, H.; Li, C.; Li, Y.; Li, B.; Zhang, Y.; Shen, S.; and Lee, Y. J. 2024 · 2024
Closest in time.
LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
Shang, Y.; Cai, M.; Xu, B.; Lee, Y. J.; and Yan, Y. 2024 · 2024
Closest in time.