Fetching the paper…
Reading the bibliography…
Current large vision-language models (VLMs) often encounter challenges such as insufficient capabilities of a single visual component and excessively long visual tokens.
Ocular dominance column development: analysis and simulation
Miller, K. D., J. B. Keller, M. P. Stryker · 1989
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S., A. Agrawal, J. Lu, et al · 2015
Earlier work this paper cites.
Modeling context in referring expressions
Yu, L., P. Poirson, S. Yang, et al · 2016
Earlier work this paper cites.
The functional diversity of retinal ganglion cells in the mouse
Baden, T., P. Berens, K. Franke, et al · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y., T. Khot, D. Summers-Stay, et al · 2017
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
Gurari, D., Q. Li, A. J. Stangl, et al · 2018
Earlier work this paper cites.
Nocaps: Novel object captioning at scale
Agrawal, H., K. Desai, Y. Wang, et al · 2019
Earlier work this paper cites.
Translating math formula images to latex sequences using deep neural networks with sequence-level training, 2019
Wang, Z., J.-C. Liu · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A., C. D. Manning · 2019
Earlier work this paper cites.
Towards vqa models that can read
Singh, A., V. Natarajan, M. Shah, et al · 2019
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., J. W. Kim, C. Hallacy, et al · 2021
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention
Touvron, H., M. Cord, M. Douze, et al · 2021
Earlier work this paper cites.
Masked autoencoders are scalable vision learners, 2021
He, K., X. Chen, S. Xie, et al · 2021
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
Caron, M., H. Touvron, I. Misra, et al · 2021
Earlier work this paper cites.
When are lemons purple? the concept association bias of clip
Yamada, Y., Y. Tang, I. Yildirim · 2022
Earlier work this paper cites.
Winoground: Probing vision and language models for visio-linguistic compositionality
Thrush, T., R. Jiang, M. Bartolo, et al · 2022
Earlier work this paper cites.
When and why vision-language models behave like bags-of-words, and what to do about it?
Yuksekgonul, M., F. Bianchi, P. Kalluri, et al · 2022
Earlier work this paper cites.
Layoutlmv3: Pre-training for document ai with unified text and image masking, 2022
Huang, Y., T. Lv, L. Cui, et al · 2022
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J., D. Li, C. Xiong, et al · 2022
Cited alongside, same era.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Lu, P., S. Mishra, T. Xia, et al · 2022
Cited alongside, same era.
Visualgpt: Data-efficient adaptation of pretrained language models for image captioning
Chen, J., H. Guo, K. Yi, et al · 2022
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., J. Donahue, P. Luc, et al · 2022
Cited alongside, same era.
The rise and potential of large language model based agents: A survey
Xi, Z., W. Chen, X. Guo, et al · 2023
Cited alongside, same era.
Visual instruction tuning
A comprehensive evaluation benchmark for multimodal large language models
Fu, C., P. Chen, Y. Shen, et al · 2023
Later among the works it cites.
Mmbench: Is your multi-modal model an all-around player?
Liu, Y., H. Duan, Y. Zhang, et al · 2023
Later among the works it cites.
Seed-bench: Benchmarking multimodal llms with generative comprehension
Li, B., R. Wang, G. Wang, et al · 2023
Later among the works it cites.
Mm-vet: Evaluating large multimodal models for integrated capabilities
Yu, W., Z. Yang, L. Li, et al · 2023
Later among the works it cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning. arxiv 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liu, H., C. Li, Q. Wu, et al · 2023
Cited alongside, same era.
Vlp: A survey on vision-language pre-training
Chen, F.-L., D.-Z. Zhang, M.-L. Han, et al · 2023
Cited alongside, same era.
Dinov2: Learning robust visual features without supervision
Oquab, M., T. Darcet, T. Moutakanni, et al · 2023
Cited alongside, same era.
What makes for good visual tokenizers for large language models?
Wang, G., Y. Ge, X. Ding, et al · 2023
Cited alongside, same era.
What’s" up" with vision-language models? investigating their struggle with spatial reasoning
Kamath, A., J. Hessel, K.-W. Chang · 2023
Cited alongside, same era.
Evaluating object hallucination in large vision-language models
Li, Y., Y. Du, K. Zhou, et al · 2023
Cited alongside, same era.
Convnext v2: Co-designing and scaling convnets with masked autoencoders
Woo, S., S. Debnath, R. Hu, et al · 2023
Cited alongside, same era.
Dai, W., J. Li, D. Li, et al · 2023
Later among the works it cites.
Shikra: Unleashing multimodal llm’s referential dialogue magic
Chen, K., Z. Zhang, W. Zeng, et al · 2023
Later among the works it cites.
Pandagpt: One model to instruction-follow them all
Su, Y., T. Lan, H. Li, et al · 2023
Later among the works it cites.
mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration, 2023
Ye, Q., H. Xu, J. Ye, et al · 2023
Later among the works it cites.
Generative multimodal models are in-context learners, 2023
Sun, Q., Y. Cui, X. Zhang, et al · 2023
Later among the works it cites.
Language is not all you need: Aligning perception with language models
Huang, S., L. Dong, W. Wang, et al · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., J. Chen, X. Shen, et al · 2023
Later among the works it cites.
Imagebind-llm: Multi-modality instruction tuning
Han, J., R. Zhang, W. Shao, et al · 2023
Later among the works it cites.
Improving compositional text-to-image generation with large vision-language models
Wen, S., G. Fang, R. Zhang, et al · 2023
Later among the works it cites.
Pointllm: Empowering large language models to understand point clouds
Xu, R., X. Wang, T. Wang, et al · 2023
Later among the works it cites.
Sharegpt4v: Improving large multi-modal models with better captions
Chen, L., J. Li, X. Dong, et al · 2023
Later among the works it cites.
To see is to believe: Prompting gpt-4v for better visual instruction tuning
Wang, J., L. Meng, Z. Weng, et al · 2023
Later among the works it cites.
Agent ai: Surveying the horizons of multimodal interaction
Durante, Z., Q. Huang, N. Wake, et al · 2024
Closest in time.