Fetching the paper…
Reading the bibliography…
High-resolution Vision-Language Models (VLMs) are widely used in multimodal tasks to enhance accuracy by preserving detailed image information.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A. 2017 · 2017
Earlier work this paper cites.
Towards VQA Models That Can Read
Singh, A.; Natarajan, V.; Shah, M.; Jiang, Y.; Chen, X.; Batra, D.; Parikh, D.; and Rohrbach, M. 2019 · 2019
Earlier work this paper cites.
Docvqa: A dataset for vqa on document images
Mathew, M.; Karatzas, D.; and Jawahar, C. 2021 · 2021
Earlier work this paper cites.
IA-REDˆ2: Interpretability-Aware Redundancy Reduction for Vision Transformers
Pan, B.; Panda, R.; Jiang, Y.; Wang, Z.; Feris, R.; and Oliva, A. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
Dynamicvit: Efficient vision transformers with dynamic token sparsification
Rao, Y.; Zhao, W.; Liu, B.; Lu, J.; Zhou, J.; and Hsieh, C.-J. 2021 · 2021
Earlier work this paper cites.
GPT3.int8(): 8-bit Matrix Multiplication for Transformers at Scale
Dettmers, T.; Lewis, M.; Belkada, Y.; and Zettlemoyer, L. 2022 · 2022
Earlier work this paper cites.
Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Lu, P.; Mishra, S.; Xia, T.; Qiu, L.; Chang, K.-W.; Zhu, S.-C.; Tafjord, O.; Clark, P.; and Kalyan, A. 2022 · 2022
Earlier work this paper cites.
ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Masry, A.; Do, X. L.; Tan, J. Q.; Joty, S.; and Hoque, E. 2022 · 2022
Earlier work this paper cites.
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Bai, J.; Bai, S.; Yang, S.; Wang, S.; Tan, S.; Wang, P.; Lin, J.; Zhou, C.; and Zhou, J. 2023 · 2023
Earlier work this paper cites.
PuMer: Pruning and Merging Tokens for Efficient Vision Language Models
Cao, Q.; Paranjape, B.; and Hajishirzi, H. 2023 · 2023
Earlier work this paper cites.
MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
Chu, X.; Qiao, L.; Lin, X.; Xu, S.; Yang, Y.; Hu, Y.; Wei, F.; Zhang, X.; Zhang, B.; Wei, X.; and Shen, C. 2023 · 2023
Earlier work this paper cites.
Evaluating Object Hallucination in Large Vision-Language Models
Li, Y.; Du, Y.; Zhou, K.; Wang, J.; Zhao, X.; and Wen, J.-R. 2023b · 2023
Cited alongside, same era.
Lin, Z.; Liu, C.; Zhang, R.; Gao, P.; Qiu, L.; Xiao, H.; Qiu, H.; Lin, C.; Shao, W.; Chen, K.; Han, J.; Huang, S.; Zhang, Y.; He, X.; Li, H.; and Qiao, Y. 2023 · 2023
Cited alongside, same era.
Visual Instruction Tuning
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023 · 2023
Cited alongside, same era.
Token pooling in vision transformers for image classification
Marin, D.; Chang, J.-H. R.; Ranjan, A.; Prabhu, A.; Rastegari, M.; and Tuzel, O. 2023 · 2023
Cited alongside, same era.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2024 · 2024
Cited alongside, same era.
Interpreting CLIP’s Image Representation via Text-Based Decomposition
Gandelsman, Y.; Efros, A. A.; and Steinhardt, J. 2024 · 2024
Closest in time.
mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
Hu, A.; Xu, H.; Ye, J.; Yan, M.; Zhang, L.; Zhang, B.; Zhang, J.; Jin, Q.; Huang, F.; and Zhou, J. 2024 · 2024
Closest in time.
FlexAttention for Efficient High-Resolution Vision-Language Models
Li, J.; Chen, D.; Cai, T.; Chen, P.; Hong, Y.; Chen, Z.; Shen, Y.; and Gan, C. 2025 · 2024
Closest in time.
Monkey: Image resolution and text label are important things for large multi-modal models
Li, Z.; Yang, B.; Liu, Q.; Ma, Z.; Zhang, S.; Yang, J.; Sun, Y.; Liu, Y.; and Bai, X. 2024 · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M.; Savinov, N.; Teplyashin, D.; Lepikhin, D.; Lillicrap, T.; Alayrac, J.-b.; Soricut, R.; Lazaridou, A.; Firat, O.; Schrittwieser, J.; et al. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cai, M.; Yang, J.; Gao, J.; and Lee, Y. J. 2024 · 2024
Cited alongside, same era.
MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer
Cao, J.; Ye, P.; Li, S.; Yu, C.; Tang, Y.; Lu, J.; and Chen, T. 2024 · 2024
Cited alongside, same era.
Honeybee: Locality-enhanced projector for multimodal llm
Cha, J.; Kang, W.; Mun, J.; and Roh, B. 2024 · 2024
Cited alongside, same era.
An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Chen, L.; Zhao, H.; Liu, T.; Bai, S.; Lin, J.; Zhou, C.; and Chang, B. 2025b · 2024
Cited alongside, same era.
Advancing High Resolution Vision-Language Models in Biomedicine
Chen, Z.; Pekis, A.; and Brown, K. 2024 · 2024
Cited alongside, same era.
Vision Transformers Need Registers
Darcet, T.; Oquab, M.; Mairal, J.; and Bojanowski, P. 2024 · 2024
Cited alongside, same era.
InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
Dong, X.; Zhang, P.; Zang, Y.; Cao, Y.; Wang, B.; Ouyang, L.; Zhang, S.; Duan, H.; Zhang, W.; Li, Y.; Yan, H.; Gao, Y.; Chen, Z.; xinyue zhang; Li, W.; Jingwen, L.; Wang, W.; Chen, K.; He, C.; ZHANG, X.; Dai, J.; Qiao, Y.; Lin, D.; and Wang, J. 2024 · 2024
Cited alongside, same era.
LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
Shang, Y.; Cai, M.; Xu, B.; Lee, Y. J.; and Yan, Y. 2024 · 2024
Closest in time.
CrossGET: Cross-Guided Ensemble of Tokens for Accelerating Vision-Language Transformers
Shi, D.; Tao, C.; Rao, A.; Yang, Z.; Yuan, C.; and Wang, J. 2024 · 2024
Closest in time.
A Simple and Effective Pruning Approach for Large Language Models
Sun, M.; Liu, Z.; Bair, A.; and Kolter, J. Z. 2024 · 2024
Closest in time.
EVIT: Event-Oriented Instruction Tuning for Event Reasoning
Tao, Z.; Chen, X.; Jin, Z.; Bai, X.; Zhao, H.; and Lou, Y. 2024 · 2024
Closest in time.
MM-LLMs: Recent Advances in MultiModal Large Language Models
Zhang, D.; Yu, Y.; Dong, J.; Li, C.; Su, D.; Chu, C.; and Yu, D. 2024a · 2024
Closest in time.
TinyLLaVA: A Framework of Small-scale Large Multimodal Models
Zhou, B.; Hu, Y.; Weng, X.; Jia, J.; Luo, J.; Liu, X.; Wu, J.; and Huang, L. 2024 · 2024
Closest in time.
LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model
Zhu, Y.; Zhu, M.; Liu, N.; Xu, Z.; and Peng, Y. 2024 · 2024
Closest in time.
MM1: methods, analysis and insights from multimodal LLM pre-training
McKinzie, B.; Gan, Z.; Fauconnier, J.-P.; Dodge, S.; Zhang, B.; Dufter, P.; Shah, D.; Du, X.; Peng, F.; Belyi, A.; et al. 2025 · 2025
Closest in time.