Fetching the paper…
Reading the bibliography…
Vision--language models (VLMs) often process visual inputs through a pretrained vision encoder, followed by a projection into the language model's embedding space via a connector component.
Analysis of a complex of statistical variables into principal components
Harold Hotelling. 1933 · 1933
Earlier work this paper cites.
Relations between two sets of variates
Harold Hotelling. 1936 · 1936
Earlier work this paper cites.
Discriminatory Analysis: Nonparametric Discrimination: Consistency Properties
E. Fix and J.L. Hodges. 1951 · 1951
Earlier work this paper cites.
Generalized procrustes analysis
John C. Gower. 1975 · 1975
Earlier work this paper cites.
Caltech-ucsd birds 200
Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie. 2011 · 2011
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll’ar, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei. 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. 2017 · 2017
Earlier work this paper cites.
Generalizing and improving bilingual word embedding mappings with a multi-step framework of linear transformations
Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2018 · 2018
Cited alongside, same era.
Perceiver: General perception with iterative attention
Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joao Carreira. 2021 · 2021
Cited alongside, same era.
Connector-s: A survey of connectors in multi-modal large language models
Xun Zhu, Zheng Zhang, Xi Chen, Yiming Shi, Miao Li, and Ji Wu. 2025 · 2021
Cited alongside, same era.
Grounding answers for visual questions asked by visually impaired people
Chongyan Chen, Samreen Anjum, and Danna Gurari. 2022 · 2022
Cited alongside, same era.
Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning
Weixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung, and James Zou. 2022 · 2022
Cited alongside, same era.
What matters when building vision-language models?
Hugo Laurençon, Léo Tronchon, Matthieu Cord, and Victor Sanh. 2024 · 2024
Later among the works it cites.
Seed-bench: Benchmarking multimodal large language models
Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang, and Ying Shan. 2024 · 2024
Later among the works it cites.
Multimodal alignment and fusion: A survey
Songtao Li and Hao Tang. 2024 · 2024
Later among the works it cites.
To preserve or to compress: An in-depth study of connector selection in multimodal large language models
Junyan Lin, Haoran Chen, Dawei Zhu, and Xiaoyu Shen. 2024 · 2024
Later among the works it cites.
Two effects, one trigger: On the modality gap, object bias, and information imbalance in contrastive vision-language representation learning
Simon Schrodi, David T Hoffmann, Max Argus, Volker Fischer, and Thomas Brox. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023 · 2023
Cited alongside, same era.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023 · 2023
Cited alongside, same era.
2024 · 2024
Cited alongside, same era.
Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models
Matt Deitke, Christopher Clark, Sangho Lee, Rohun Tripathi, Yue Yang, Jae Sung Park, Mohammadreza Salehi, Niklas Muennighoff, Kyle Lo, Luca Soldaini, et al. 2024 · 2024
Cited alongside, same era.
Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou. 2024 · 2024
Cited alongside, same era.
Lion: Empowering multimodal large language model with dual-level visual knowledge
Gongwei Chen, Leyang Shen, Rui Shao, Xiang Deng, and Liqiang Nie. 2024a
Cited in the paper.
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. 2024b
Cited in the paper.
Later among the works it cites.
Generative multimodal models are in-context learners
Quan Sun, Yufeng Cui, Xiaosong Zhang, Fan Zhang, Qiying Yu, Yueze Wang, Yongming Rao, Jingjing Liu, Tiejun Huang, and Xinlong Wang. 2024 · 2024
Later among the works it cites.
Eyes wide shut? exploring the visual shortcomings of multimodal llms
Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann Lecun, and Saining Xie. 2024 · 2024
Later among the works it cites.
Why are visually-grounded language models bad at image classification?
Yuhui Zhang, Alyssa Unell, Xiaohan Wang, Dhruba Ghosh, Yuchang Su, Ludwig Schmidt, and Serena Yeung-Levy. 2024 · 2024
Later among the works it cites.
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. 2025 · 2025
Closest in time.
MM1.5: Methods, analysis & insights from multimodal LLM fine-tuning
Haotian Zhang, Mingfei Gao, Zhe Gan, Philipp Dufter, Nina Wenzel, Forrest Huang, Dhruti Shah, Xianzhi Du, Bowen Zhang, Yanghao Li, Sam Dodge, Keen You, Zhen Yang, Aleksei Timofeev, Mingze Xu, Hong-You Chen, Jean-Philippe Fauconnier, Zhengfeng Lai, Haoxuan You, Zirui Wang, Afshin Dehghan, Peter Grasch, and Yinfei Yang. 2025 · 2025
Closest in time.