Fetching the paper…
Reading the bibliography…
Recent Multimodal Large Language Models(MLLMs) often use a large number of visual tokens to compensate their visual shortcoming, leading to excessive computation and obvious visual redundancy.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017 · 2017
Earlier work this paper cites.
Vaswani, A. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A.; and Manning, C. D. 2019 · 2019
Earlier work this paper cites.
Towards vqa models that can read
Singh, A.; Natarajan, V.; Shah, M.; Jiang, Y.; Chen, X.; Batra, D.; Parikh, D.; and Rohrbach, M. 2019 · 2019
Earlier work this paper cites.
DocVQA: A Dataset for VQA on Document Images
Mathew, M.; Karatzas, D.; Manmatha, R.; and Jawahar, C. V. 2020 · 2021
Earlier work this paper cites.
IA-RED2: Interpretability-Aware Redundancy Reduction for Vision Transformers
Pan, B.; Panda, R.; Jiang, Y.; Wang, Z.; Feris, R.; and Oliva, A. 2021 · 2021
Earlier work this paper cites.
Dynamicvit: Efficient vision transformers with dynamic token sparsification
Rao, Y.; Zhao, W.; Liu, B.; Lu, J.; Zhou, J.; and Hsieh, C.-J. 2021 · 2021
Earlier work this paper cites.
Token merging: Your vit but faster
Bolya, D.; Fu, C.-Y.; Dai, X.; Zhang, P.; Feichtenhofer, C.; and Hoffman, J. 2022 · 2022
Earlier work this paper cites.
Not all patches are what you need: Expediting vision transformers via token reorganizations
Liang, Y.; Ge, C.; Tong, Z.; Song, Y.; Wang, J.; and Xie, P. 2022 · 2022
Cited alongside, same era.
Norm-based noisy corpora filtering and refurbishing in neural machine translation
Lu, Y.; and Zhang, J. 2022 · 2022
Cited alongside, same era.
Chartqa: A benchmark for question answering about charts with visual and logical reasoning
Masry, A.; Long, D. X.; Tan, J. Q.; Joty, S.; and Hoque, E. 2022 · 2022
Cited alongside, same era.
Evo-vit: Slow-fast token evolution for dynamic vision transformer
Xu, Y.; Zhang, Z.; Zhang, M.; Sheng, K.; Li, K.; Dong, W.; Zhang, L.; Xu, C.; and Sun, X. 2022 · 2022
Cited alongside, same era.
Focal modulation networks
Yang, J.; Li, C.; Dai, X.; and Gao, J. 2022 · 2022
Cited alongside, same era.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023 · 2023
Later among the works it cites.
MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer
Cao, J.; Ye, P.; Li, S.; Yu, C.; Tang, Y.; Lu, J.; and Chen, T. 2024 · 2024
Later among the works it cites.
DeepSeek-Coder: When the Large Language Model Meets Programming–The Rise of Code Intelligence
Guo, D.; Zhu, Q.; Yang, D.; Xie, Z.; Dong, K.; Zhang, W.; Chen, G.; Bi, X.; Wu, Y.; Li, Y.; et al. 2024 · 2024
Later among the works it cites.
Lmms-eval: Accelerating the development of large multimoal models
Li, B.; Zhang, P.; Zhang, K.; Pu, F.; Du, X.; Dong, Y.; Liu, H.; Zhang, Y.; Zhang, G.; Li, C.; et al. 2024 · 2024
Later among the works it cites.
Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Cited alongside, same era.
PuMer: Pruning and merging tokens for efficient vision language models
Cao, Q.; Paranjape, B.; and Hajishirzi, H. 2023 · 2023
Cited alongside, same era.
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Fu, C.; Chen, P.; Shen, Y.; Qin, Y.; Zhang, M.; Lin, X.; Yang, J.; Zheng, X.; Li, K.; Sun, X.; Wu, Y.; and Ji, R. 2023 · 2023
Cited alongside, same era.
Mmbench: Is your multi-modal model an all-around player?
Liu, Y.; Duan, H.; Zhang, Y.; Li, B.; Zhang, S.; Zhao, W.; Yuan, Y.; Wang, J.; He, C.; Liu, Z.; et al. 2023 · 2023
Cited alongside, same era.
Beyond attentive tokens: Incorporating token importance and diversity for efficient vision transformers
Long, S.; Zhao, Z.; Pi, J.; Wang, S.; and Wang, J. 2023 · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Team, G.; Anil, R.; Borgeaud, S.; Wu, Y.; Alayrac, J.-B.; Yu, J.; Soricut, R.; Schalkwyk, J.; Dai, A. M.; Hauth, A.; et al. 2023 · 2023
Cited alongside, same era.
Bai, J.; Bai, S.; Chu, Y.; Cui, Z.; Dang, K.; Deng, X.; Fan, Y.; Ge, W.; Han, Y.; Huang, F.; et al. 2023a
Cited in the paper.
Lin, Z.; Lin, M.; Lin, L.; and Ji, R. 2024 · 2024
Later among the works it cites.
Cheap and quick: Efficient vision-language instruction tuning for large language models
Luo, G.; Zhou, Y.; Ren, T.; Chen, S.; Sun, X.; and Ji, R. 2024 · 2024
Later among the works it cites.
Introducing meta llama 3: The most capable openly available llm to date
Meta, A. 2024 · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M.; Savinov, N.; Teplyashin, D.; Lepikhin, D.; Lillicrap, T.; Alayrac, J.-b.; Soricut, R.; Lazaridou, A.; Firat, O.; Schrittwieser, J.; et al. 2024 · 2024
Later among the works it cites.
Zero-TPrune: Zero-shot token pruning through leveraging of the attention graph in pre-trained transformers
Wang, H.; Dedhia, B.; and Jha, N. K. 2024 · 2024
Later among the works it cites.
DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
Yao, L.; Li, L.; Ren, S.; Wang, L.; Liu, Y.; Sun, X.; and Hou, L. 2024 · 2024
Later among the works it cites.
Zhang, P.; Dong, X.; Zang, Y.; Cao, Y.; Qian, R.; Chen, L.; Guo, Q.; Duan, H.; Wang, B.; Ouyang, L.; Zhang, S.; Zhang, W.; Li, Y.; Gao, Y.; Sun, P.; Zhang, X.; Li, W.; Li, J.; Wang, W.; Yan, H.; He, C.; Zhang, X.; Chen, K.; Dai, J.; Qiao, Y.; Lin, D.; and Wang, J. 2024 · 2024
Later among the works it cites.