Fetching the paper…
Reading the bibliography…
Transforming a large language model (LLM) into a vision-language model (VLM) can be achieved by mapping the visual tokens from a vision encoder into the embedding space of an LLM.
The symbol grounding problem
Harnad, S · 1990
Earlier work this paper cites.
Visual Genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.-J., Shamma, D. A., et al · 2017
Earlier work this paper cites.
Deep contextualized word representations
Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L · 2018
Earlier work this paper cites.
What does this word mean? explaining contextualized embeddings with natural language definition
Chang, T.-Y. and Chen, Y.-N · 2019
Earlier work this paper cites.
How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings
Ethayarajh, K · 2019
Earlier work this paper cites.
The bottom-up evolution of representations in the transformer: A study with machine translation and language modeling objectives
Voita, E., Sennrich, R., and Titov, I · 2019
Earlier work this paper cites.
Does BERT make any sense? interpretable word sense disambiguation with contextualized embeddings
Wiedemann, G., Remus, S., Chawla, A., and Biemann, C · 2019
Earlier work this paper cites.
VisBERT: Hidden-state visualizations for transformers
Aken, B. v., Winter, B., Löser, A., and Gers, F. A · 2020
Earlier work this paper cites.
Experience grounds language
Bisk, Y., Holtzman, A., Thomason, J., Andreas, J., Bengio, Y., Chai, J., Lapata, M., Lazaridou, A., May, J., Nisnevich, A., Pinto, N., and Turian, J · 2020
Earlier work this paper cites.
spaCy: Industrial-strength Natural Language Processing in Python
Honnibal, M., Montani, I., Van Landeghem, S., and Boyd, A · 2020
Earlier work this paper cites.
interpreting gpt: the logit lens, 2020
nostalgebraist · 2020
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Geva, M., Schuster, R., Berant, J., and Levy, O · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Earlier work this paper cites.
ClipCap: CLIP Prefix for Image Captioning
Mokady, R., Hertz, A., and Bermano, A. H · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Earlier work this paper cites.
All bark and no bite: Rogue dimensions in transformer language models obscure representational quality
Timkey, W. and van Schijndel, M · 2021
Earlier work this paper cites.
Multimodal few-shot learning with frozen language models
Tsimpoukelli, M., Menick, J., Cabi, S., Eslami, S. M. A., Vinyals, O., and Hill, F · 2021
Earlier work this paper cites.
Probing classifiers: Promises, shortcomings, and advances
Belinkov, Y · 2022
Earlier work this paper cites.
Large scale substitution-based word sense induction
Eyal, M., Sadde, S., Taub-Tabib, H., and Goldberg, Y · 2022
Earlier work this paper cites.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Geva, M., Caciularu, A., Wang, K., and Goldberg, Y · 2022
Earlier work this paper cites.
Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning
Liang, W., Zhang, Y., Kwon, Y., Yeung, S., and Zou, J · 2022
Earlier work this paper cites.
Frozen pretrained transformers as universal computation engines
Lu, K., Grover, A., Abbeel, P., and Mordatch, I · 2022
Earlier work this paper cites.
Mapping language models to grounded conceptual spaces
Patel, R. and Pavlick, E · 2022
Earlier work this paper cites.
Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., et al · 2023
Earlier work this paper cites.
Soft prompting might be a bug, not a feature
Bailey, L., Ahdritz, G., Kleiman, A., Swaroop, S., Doshi-Velez, F., and Pan, W · 2023
Earlier work this paper cites.
Eliciting Latent Predictions from Transformers with the Tuned Lens
Belrose, N., Furman, Z., Smith, L., Halawi, D., Ostrovsky, I., McKinney, L., Biderman, S., and Steinhardt, J · 2023
Earlier work this paper cites.
Dissecting recall of factual associations in auto-regressive language models
Geva, M., Bastings, J., Filippova, K., and Globerson, A · 2023
Cited alongside, same era.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2023
Cited alongside, same era.
MAPL: Parameter-efficient adaptation of unimodal pre-trained models for vision-language few-shot prompting
Mañas, O., Rodriguez Lopez, P., Ahmadi, S., Nematzadeh, A., Goyal, Y., and Agrawal, A · 2023
Cited alongside, same era.
Linearly mapping from image to text space
Merullo, J., Castricato, L., Eickhoff, C., and Pavlick, E · 2023
Cited alongside, same era.
Cross-modal fine-tuning: Align then refine
Shen, J., Li, L., Dery, L. M., Staten, C., Khodak, M., Neubig, G., and Talwalkar, A · 2023
Cited alongside, same era.
Yang, A., Yang, B., Hui, B., Zheng, B., Yu, B., Zhou, C., Li, C., Li, C., Liu, D., Huang, F., Dong, G., Wei, H., Lin, H., Tang, J., Wang, J., Yang, J., Tu, J., Zhang, J., Ma, J., Xu, J., Zhou, J., Bai, J., He, J., Lin, J., Dang, K., Lu, K., Chen, K.-Y., Yang, K., Li, M., Xue, M., Ni, N., Zhang, P., Wang, P., Peng, R., Men, R., Gao, R., Lin, R., Wang, S., Bai, S., Tan, S., Zhu, T., Li, T., Liu, T., Ge, W., Deng, X., Zhou, X., Ren, X., Zhang, X., Wei, X., Ren, X., Fan, Y., Yao, Y., Zhang, Y., Wan, Y., Chu, Y., Cui, Z., Zhang, Z., and Fan, Z.-W · 2024
Later among the works it cites.
Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., et al · 2025
Later among the works it cites.
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Deitke, M., Clark, C., Lee, S., Tripathi, R., Yang, Y., Park, J. S., Salehi, M., Muennighoff, N., Lo, K., Soldaini, L., Lu, J., Anderson, T., Bransom, E., Ehsani, K., Ngo, H., Chen, Y., Patel, A., Yatskar, M., Callison-Burch, C., Head, A., Hendrix, R., Bastani, F., VanderBilt, E., Lambert, N., Chou, Y., Chheda, A., Sparks, J., Skjonsberg, S., Schmitz, M., Sarnat, A., Bischoff, B., Walsh, P., Newell, C., Wolters, P., Gupta, T., Zeng, K.-H., Borchardt, J., Groeneveld, D., Nam, C., Lebrecht, S., Wittlif, C., Schoenick, C., Michel, O., Krishna, R., Weihs, L., Smith, N. A., Hajishirzi, H., Girshick, R., Farhadi, A., and Kembhavi, A · 2025
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Team, G., Anil, R., Borgeaud, S., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., Millican, K., et al · 2023
Cited alongside, same era.
Sigmoid loss for language image pre-training
Zhai, X., Mustafa, B., Kolesnikov, A., and Beyer, L · 2023
Cited alongside, same era.
PaliGemma: A versatile 3B VLM for transfer
Beyer, L., Steiner, A., Pinto, A. S., Kolesnikov, A., Wang, X., Salz, D., Neumann, M., Alabdulmohsin, I., Tschannen, M., Bugliarello, E., et al · 2024
Cited alongside, same era.
Analyzing the language of visual tokens
Chan, D. M., Corona, R., Park, J., Cho, C. J., Bai, Y., and Darrell, T · 2024
Cited alongside, same era.
INViTE: INterpret and control vision-language models with text explanations
Chen, H., Yang, J., Vondrick, C., and Mao, C · 2024
Cited alongside, same era.
Interpreting CLIP’s image representation via text-based decomposition
Gandelsman, Y., Efros, A. A., and Steinhardt, J · 2024
Cited alongside, same era.
What explains the success of cross-modal fine-tuning with ORCA?
García-de Herreros, P., Gautam, V., Slusallek, P., Klakow, D., and Mosbach, M · 2024
Cited alongside, same era.
How do multilingual language models remember facts?
Fierro, C., Foroutan, N., Elliott, D., and Søgaard, A · 2025
Later among the works it cites.
Hidden in plain sight: VLMs overlook their visual representations
Fu, S., Bonnen, T., Guillory, D., and Darrell, T · 2025
Later among the works it cites.
Training large language models to reason in a continuous latent space
Hao, S., Sukhbaatar, S., Su, D., Li, X., Hu, Z., Weston, J. E., and Tian, Y · 2025
Later among the works it cites.
LLaVA-onevision: Easy visual task transfer
Li, B., Zhang, Y., Guo, D., Zhang, R., Li, F., Zhang, H., Zhang, K., Zhang, P., Li, Y., Liu, Z., and Li, C · 2025
Later among the works it cites.
LangBridge: Interpreting Image as a Combination of Language Embeddings
Liao, J., Niu, Y., Meng, F., Li, H., Tian, C., Du, Y., Xiong, Y., Li, D., Zhu, X., Yuan, L., Dai, J., and Cheng, Y · 2025
Later among the works it cites.
Alignvlm: Bridging vision and language latent spaces for multimodal document understanding
Masry, A., Rodriguez, J. A., Zhang, T., Wang, S., Wang, C., Feizi, A., Suresh, A. K., Puri, A., Jian, X., Noel, P.-A., Madhusudhan, S. T., Pedersoli, M., Liu, B., Chapados, N., Bengio, Y., Hoque, E., Pal, C., Laradji, I. H., Vazquez, D., Taslakian, P., Gella, S., and Rajeswar, S · 2025
Later among the works it cites.
Towards interpreting visual information processing in vision-language models
Neo, C., Ong, L., Torr, P., Geva, M., Krueger, D., and Barez, F · 2025
Later among the works it cites.
Ògúnrèmí, T., Manning, C. D., Jurafsky, D., and Livescu, K · 2025
Later among the works it cites.
Seeing what tastes good: Revisiting multimodal distributional semantics in the billion parameter era
Oneata, D., Elliott, D., and Frank, S · 2025
Later among the works it cites.
Interpreting the linear structure of vision-language model embedding spaces
Papadimitriou, I., Su, H., Fel, T., Kakade, S. M., and Gil, S · 2025
Later among the works it cites.
GLSim: Detecting object hallucinations in LVLMs via global-local similarity
Park, S. and Li, S · 2025
Later among the works it cites.
Beyond logit lens: Contextual embeddings for robust hallucination detection & grounding in VLMs
Phukan, A., Divyansh, Morj, H. K., Vaishnavi, Saxena, A., and Goswami, K · 2025
Later among the works it cites.
How visual representations map to language feature space in multimodal llms
Venhoff, C., Khakzar, A., Joseph, S., Torr, P., and Nanda, N · 2025
Later among the works it cites.
The semantic hub hypothesis: Language models share semantic representations across languages and modalities
Wu, Z., Yu, X. V., Yogatama, D., Lu, J., and Kim, Y · 2025
Later among the works it cites.
Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., et al · 2025
Later among the works it cites.
Painting with words: Elevating detailed image captioning with benchmark and alignment learning
Ye, Q., Zeng, X., Li, F., Li, C., and Fan, H · 2025
Later among the works it cites.
Cross-modal information flow in multimodal large language models
Zhang, Z., Yadav, S., Han, F., and Shutova, E · 2025
Later among the works it cites.
Transfusion: Predict the next token and diffuse images with one multi-modal model
Zhou, C., YU, L., Babu, A., Tirumala, K., Yasunaga, M., Shamis, L., Kahn, J., Ma, X., Zettlemoyer, L., and Levy, O · 2025
Later among the works it cites.
PerceptionLM: Open-access data and models for detailed visual understanding
Cho, J. H., Madotto, A., Mavroudi, E., Afouras, T., Nagarajan, T., Maaz, M., Song, Y., Ma, T., Hu, S., Jain, S., Martin, M., Wang, H., Rasheed, H. A., Sun, P., Huang, P.-Y., Bolya, D., Ravi, N., Jain, S., Stark, T., Moon, S., Damavandi, B., Lee, V., Westbury, A., Khan, S., Kraehenbuehl, P., Dollar, P., Torresani, L., Grauman, K., and Feichtenhofer, C · 2026
Closest in time.
Learning to see before seeing: Demystifying LLM visual priors from language pre-training
Han, J., Tong, S., Fan, D., Ren, Y., Sinha, K., Torr, P., and Kokkinos, F · 2026
Closest in time.
Same task, different circuits: Disentangling modality-specific mechanisms in VLMs
Nikankin, Y., Arad, D., Gandelsman, Y., and Belinkov, Y · 2026
Closest in time.
Representations of text and images align from layer one
Wybitul, E., Rando, J., Tramèr, F., and Fort, S · 2026
Closest in time.