Fetching the paper…
Reading the bibliography…
The established redundancy in visual tokens within large vision-language models allows pruning to effectively reduce their substantial computational demands.
A mathematical theory of communication
Claude E Shannon · 1948
Earlier work this paper cites.
The merging of the senses
Barry E Stein and M Alex Meredith · 1993
Earlier work this paper cites.
Integration of visual and linguistic information in spoken language comprehension
Michael K Tanenhaus, Michael J Spivey-Knowlton, Kathleen M Eberhard, and Julie C Sedivy · 1995
Earlier work this paper cites.
Eye movements in reading and information processing: 20 years of research
Keith Rayner · 1998
Earlier work this paper cites.
High-level scene perception
John M Henderson and Andrew Hollingworth · 1999
Earlier work this paper cites.
Incremental interpretation at verbs: Restricting the domain of subsequent reference
Gerry TM Altmann and Yuki Kamide · 1999
Earlier work this paper cites.
Analyzing ‘visual world’eyetracking data using multilevel logistic regression
Dale J Barr · 2008
Earlier work this paper cites.
Using the visual world paradigm to study language processing: A review and critical evaluation
Falk Huettig, Joost Rommers, and Antje S Meyer · 2011
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick · 2015
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning · 2019
Earlier work this paper cites.
Towards vqa models that can read
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach · 2019
Earlier work this paper cites.
Resolving multisensory and attentional influences across cortical depth in sensory cortices
Remi Gau, Pierre-Louis Bazin, Robert Trampel, Robert Turner, and Uta Noppeney · 2020
Earlier work this paper cites.
A general survey on attention mechanisms in deep learning
Gianni Brauwers and Flavius Frasincar · 2021
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Earlier work this paper cites.
Git: A generative image-to-text transformer for vision and language
Jianfeng Wang, Zhengyuan Yang, Xiaowei Hu, Linjie Li, Kevin Lin, Zhe Gan, Zicheng Liu, Ce Liu, and Lijuan Wang · 2022
Earlier work this paper cites.
Blenderbot 3: a deployed conversational agent that continually learns to responsibly engage
Kurt Shuster, Jing Xu, Mojtaba Komeili, Da Ju, Eric Michael Smith, Stephen Roller, Megan Ung, Moya Chen, Kushal Arora, Joshua Lane, et al · 2022
Cited alongside, same era.
Token merging: Your vit but faster
Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman · 2022
Cited alongside, same era.
Adaptive sparse vit: Towards learnable adaptive token pruning by fully exploiting self-attention
Xiangcheng Liu, Tianyi Wu, and Guodong Guo · 2022
Cited alongside, same era.
Vision-language pre-training with triple contrastive learning
Jinyu Yang, Jiali Duan, Son Tran, Yi Xu, Sampath Chanda, Liqun Chen, Belinda Zeng, Trishul Chilimbi, and Junzhou Huang · 2022
Cited alongside, same era.
Scienceqa: A novel resource for question answering on scholarly articles
Visionzip: Longer is better but not necessary in vision language models
Senqiao Yang, Yukang Chen, Zhuotao Tian, Chengyao Wang, Jingyao Li, Bei Yu, and Jiaya Jia · 2024
Later among the works it cites.
An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Liang Chen, Haozhe Zhao, Tianyu Liu, Shuai Bai, Junyang Lin, Chang Zhou, and Baobao Chang · 2024
Later among the works it cites.
[cls] attention is all you need for training-free visual token pruning: Make vlm inference faster
Qizhe Zhang, Aosong Cheng, Ming Lu, Zhiyong Zhuo, Minqi Wang, Jiajun Cao, Shaobo Guo, Qi She, and Shanghang Zhang · 2024
Later among the works it cites.
Yefei He, Feng Chen, Jing Liu, Wenqi Shao, Hong Zhou, Kaipeng Zhang, and Bohan Zhuang · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tanik Saikh, Tirthankar Ghosal, Amish Mittal, Asif Ekbal, and Pushpak Bhattacharyya · 2022
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Cited alongside, same era.
mplug-owl: Modularization empowers large language models with multimodality
Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, et al · 2023
Cited alongside, same era.
Llama-vid: An image is worth 2 tokens in large language models
Yanwei Li et al · 2023
Cited alongside, same era.
Haotian Liu, Pengfei Zhang, Yizhu Xu, Hang Zhang, Xin Li, Lidong Bing, et al · 2023
Cited alongside, same era.
Dynamic token pruning in plain vision transformers for semantic segmentation
Quan Tang, Bowen Zhang, Jiajun Liu, Fagui Liu, and Yifan Liu · 2023
Cited alongside, same era.
Evaluating object hallucination in large vision-language models
Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen · 2023
Cited alongside, same era.
Openvla: An open-source vision-language-action model
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Paul Foster, Grace Lam, Pannag Sanketi, et al · 2024
Cited alongside, same era.
Revealing vision-language integration in the brain with multimodal networks
Vighnesh Subramaniam, Colin Conwell, Christopher Wang, Gabriel Kreiman, Boris Katz, Ignacio Cases, and Andrei Barbu · 2024
Later among the works it cites.
Llavanext: Improved reasoning, ocr, and world knowledge, 2024
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee · 2024
Later among the works it cites.
Senna: Bridging large vision-language models and end-to-end autonomous driving
Bo Jiang, Shaoyu Chen, Bencheng Liao, Xingyu Zhang, Wei Yin, Qian Zhang, Chang Huang, Wenyu Liu, and Xinggang Wang · 2024
Later among the works it cites.
Video token sparsification for efficient multimodal llms in autonomous driving
Yunsheng Ma, Amr Abdelraouf, Rohit Gupta, Ziran Wang, and Kyungtae Han · 2024
Later among the works it cites.
B-vllm: A vision large language model with balanced spatio-temporal tokens
Zhuqiang Lu, Zhenfei Yin, Mengwei He, Zhihui Wang, Zicheng Liu, Zhiyong Wang, and Kun Hu · 2024
Later among the works it cites.
Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024
Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, Yunsheng Wu, and Rongrong Ji · 2024
Later among the works it cites.
Mmbench: Is your multi-modal model an all-around player?, 2024
Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, Kai Chen, and Dahua Lin · 2024
Later among the works it cites.
[cls] attention is all you need for training-free visual token pruning: Make vlm inference faster
Qizhe Zhang, Aosong Cheng, Ming Lu, Zhiyong Zhuo, MinQi Wang, Jiajun Cao, Shaobo Guo, Qi She, and Shanghang Zhang · 2024
Later among the works it cites.
Vqa 2 : Visual question answering for video quality assessment
Ziheng Jia, Zicheng Zhang, Jiaying Qian, Haoning Wu, Wei Sun, Chunyi Li, Xiaohong Liu, Weisi Lin, Guangtao Zhai, and Xiongkuo Min · 2024
Later among the works it cites.
Sparsevlm: Visual token sparsification for efficient vision-language model inference
Yuan Zhang, Chun-Kai Fan, Junpeng Ma, Wenzhao Zheng, Tao Huang, Kuan Cheng, Denis Gudovskiy, Tomoyuki Okuno, Yohei Nakata, Kurt Keutzer, et al · 2025
Closest in time.
Hired: Attention-guided token dropping for efficient inference of high-resolution vision-language models
Kazi Hasan Ibn Arif, JinYi Yoon, Dimitrios S Nikolopoulos, Hans Vandierendonck, Deepu John, and Bo Ji · 2025
Closest in time.
Vltp: Vision-language guided token pruning for task-oriented segmentation
Hanning Chen, Yang Ni, Wenjun Huang, Yezi Liu, SungHeon Jeong, Fei Wen, Nathaniel D Bastian, Hugo Latapie, and Mohsen Imani · 2025
Closest in time.
Cheng Yang, Yang Sui, Jinqi Xiao, Lingyi Huang, Yu Gong, Chendi Li, Jinghua Yan, Yu Bai, Ponnuswamy Sadayappan, Xia Hu, et al · 2025
Closest in time.