Fetching the paper…
Reading the bibliography…
We introduce Lavender, a simple supervised fine-tuning (SFT) method that boosts the performance of advanced vision-language models (VLMs) by leveraging state-of-the-art image generation models such as Stable Diffusion.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P., Lai, A., Hodosh, M., and Hockenmaier, J · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Chen, X., Fang, H., Lin, T.-Y., Vedantam, R., Gupta, S., Dollár, P., and Zitnick, C. L · 2015
Earlier work this paper cites.
A diagram is worth a dozen images
Kembhavi, A., Salvato, M., Kolve, E., Seo, M., Hajishirzi, H., and Farhadi, A · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
Gurari, D., Li, Q., Stangl, A. J., Guo, A., Lin, C., Grauman, K., Luo, J., and Bigham, J. P · 2018
Earlier work this paper cites.
Squeeze-and-excitation networks
Hu, J., Shen, L., and Sun, G · 2018
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Marino, K., Rastegari, M., Farhadi, A., and Mottaghi, R · 2019
Earlier work this paper cites.
Ocr-vqa: Visual question answering by reading text in images
Mishra, A., Shekhar, S., Singh, A. K., and Chakraborty, A · 2019
Earlier work this paper cites.
Towards VQA models that can read
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M · 2019
Earlier work this paper cites.
Quantifying attention flow in transformers
Abnar, S. and Zuidema, W · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A · 2020
Earlier work this paper cites.
The hateful memes challenge: Detecting hate speech in multimodal memes
Kiela, D., Firooz, H., Mohan, A., Goswami, V., Singh, A., Ringshia, P., and Testuggine, D · 2020
Earlier work this paper cites.
Perceiver: General perception with iterative attention
Jaegle, A., Gimeno, F., Brock, A., Vinyals, O., Zisserman, A., and Carreira, J · 2021
Earlier work this paper cites.
Highly accurate protein structure prediction with alphafold
Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Žídek, A., Potapenko, A., et al · 2021
Earlier work this paper cites.
DocVQA: A dataset for vqa on document images
Mathew, M., Karatzas, D., and Jawahar, C. V · 2021
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models, 2021
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Earlier work this paper cites.
Prompt-to-prompt image editing with cross attention control
Hertz, A., Mokady, R., Tenenbaum, J., Aberman, K., Pritch, Y., and Cohen-Or, D · 2022
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J., Li, D., Xiong, C., and Hoi, S · 2022
Earlier work this paper cites.
ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Masry, A., Long, D. X., Tan, J. Q., Joty, S., and Hoque, E · 2022
Earlier work this paper cites.
Infographicvqa
Mathew, M., Bagal, V., Tito, R., Karatzas, D., Valveny, E., and Jawahar, C · 2022
Earlier work this paper cites.
Null-text inversion for editing real images using guided diffusion models, 2022
Mokady, R., Hertz, A., Aberman, K., Pritch, Y., and Cohen-Or, D · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
ScienceQA: A novel resource for question answering on scholarly articles
Saikh, T., Ghosal, T., Mittal, A., Ekbal, A., and Bhattacharyya, P · 2022
Cited alongside, same era.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al · 2022
Cited alongside, same era.
Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., Hui, B., Ji, L., Li, M., Lin, J., Lin, R., Liu, D., Liu, G., Lu, C., Lu, K., Ma, J., Men, R., Ren, X., Ren, X., Tan, C., Tan, S., Tu, J., Wang, P., Wang, S., Wang, W., Wu, S., Xu, B., Xu, J., Yang, A., Yang, H., Yang, J., Yang, S., Yao, Y., Yu, B., Yuan, H., Yuan, Z., Zhang, J., Zhang, X., Zhang, Y., Zhang, Z., Zhou, C., Zhou, J., Zhou, X., and Zhu, T · 2023
Cited alongside, same era.
Tem-adapter: Adapting image-text pretraining for video question answer
Vlmevalkit: An open-source toolkit for evaluating large multi-modality models
Duan, H., Yang, J., Qiao, Y., Fang, X., Chen, L., Liu, Y., Dong, X., Zang, Y., Zhang, P., Wang, J., et al · 2024
Later among the works it cites.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Later among the works it cites.
Generative visual instruction tuning
Hernandez, J., Villegas, R., and Ordonez, V · 2024
Later among the works it cites.
Segment and caption anything
Huang, X., Wang, J., Tang, Y., Zhang, Z., Hu, H., Lu, J., Wang, L., and Liu, Z · 2024
Later among the works it cites.
Prismatic vlms: Investigating the design space of visually-conditioned language models
Karamcheti, S., Nair, S., Balakrishna, A., Liang, P., Kollar, T., and Sadigh, D · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, G., Liu, X., Wang, G., Zhang, K., Torr, P. H., Zhang, X.-P., and Tang, Y · 2023
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2023
Cited alongside, same era.
Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S · 2023
Cited alongside, same era.
MME: A comprehensive evaluation benchmark for multimodal large language models
Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Qiu, Z., Lin, W., Yang, J., Zheng, X., et al · 2023
Cited alongside, same era.
Llama-adapter v2: Parameter-efficient visual instruction model
Gao, P., Han, J., Zhang, R., Lin, Z., Geng, S., Zhou, A., Zhang, W., Lu, P., He, C., Yue, X., et al · 2023
Cited alongside, same era.
From clip to dino: Visual encoders shout in multi-modal large language models
Jiang, D., Liu, Y., Liu, S., Zhao, J., Zhang, H., Gao, Z., Zhang, X., Li, J., and Xiong, H · 2023
Cited alongside, same era.
An image is worth multiple words: Learning object level concepts using multi-concept prompt learning
Jin, C., Tanno, R., Saseendran, A., Diethe, T., and Teare, P · 2023
Cited alongside, same era.
The flan collection: Designing data and methods for effective instruction tuning
Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., et al · 2023
Cited alongside, same era.
Later among the works it cites.
Generating images with multimodal language models
Koh, J. Y., Fried, D., and Salakhutdinov, R. R · 2024
Later among the works it cites.
Llava-onevision: Easy visual task transfer
Li, B., Zhang, Y., Guo, D., Zhang, R., Li, F., Zhang, H., Zhang, K., Zhang, P., Li, Y., Liu, Z., et al · 2024
Later among the works it cites.
Training-free open-ended object detection and segmentation via attention as prompts
Lin, Z., Wang, Y., and Tang, Z · 2024
Later among the works it cites.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024b
Liu, H., Li, C., Li, Y., Li, B., Zhang, Y., Shen, S., and Lee, Y. J · 2024
Later among the works it cites.
Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action
Lu, J., Clark, C., Lee, S., Zhang, Z., Khosla, S., Marten, R., Hoiem, D., and Kembhavi, A · 2024
Later among the works it cites.
Matos, J., Chen, S., Placino, S., Li, Y., Pardo, J. C. C., Idan, D., Tohyama, T., Restrepo, D., Nakayama, L. F., Pascual-Leone, J. M., et al · 2024
Later among the works it cites.
GPT-4o mini system card, 2024
OpenAI · 2024
Later among the works it cites.
Eagle: Exploring the design space for multimodal llms with mixture of encoders
Shi, M., Liu, F., Wang, S., Liao, S., Radhakrishnan, S., Huang, D.-A., Yin, H., Sapra, K., Yacoob, Y., Shi, H., et al · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Team, G · 2024
Later among the works it cites.
Towards generalist biomedical ai
Tu, T., Azizi, S., Driess, D., Schaekermann, M., Amin, M., Chang, P.-C., Carroll, A., Lau, C., Tanno, R., Ktena, I., et al · 2024
Later among the works it cites.
Advancing multimodal medical capabilities of gemini
Yang, L., Xu, S., Sellergren, A., Kohlberger, T., Zhou, Y., Ktena, I., Kiraly, A., Ahmed, F., Hormozdiari, F., Jaroensri, T., et al · 2024
Later among the works it cites.
Minicpm-v: A gpt-4v level mllm on your phone
Yao, Y., Yu, T., Zhang, A., Wang, C., Cui, J., Zhu, H., Cai, T., Li, H., Zhao, W., He, Z., Chen, Q., Zhou, H., Zou, Z., Zhang, H., Hu, S., Zheng, Z., Zhou, J., Cai, J., Han, X., Zeng, G., Li, D., Liu, Z., and Sun, M · 2024
Later among the works it cites.
Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness
Yu, T., Zhang, H., Yao, Y., Dang, Y., Chen, D., Lu, X., Cui, G., He, T., Liu, Z., Chua, T.-S., and Sun, M · 2024
Later among the works it cites.
MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI
Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y., et al · 2024
Later among the works it cites.
Transfusion: Predict the next token and diffuse images with one multi-modal model
Zhou, C., Yu, L., Babu, A., Tirumala, K., Yasunaga, M., Shamis, L., Kahn, J., Ma, X., Zettlemoyer, L., and Levy, O · 2024
Later among the works it cites.
Multimodal c4: An open, billion-scale corpus of images interleaved with text
Zhu, W., Hessel, J., Awadalla, A., Gadre, S. Y., Dodge, J., Fang, A., Yu, Y., Schmidt, L., Wang, W. Y., and Choi, Y · 2024
Later among the works it cites.
Mova: Adapting mixture of vision experts to multimodal context
Zong, Z., Ma, B., Shen, D., Song, G., Shao, H., Jiang, D., Li, H., and Liu, Y · 2024
Later among the works it cites.
Brave: Broadening the visual encoding of vision-language models
Kar, O. F., Tonioni, A., Poklukar, P., Kulshrestha, A., Zamir, A., and Tombari, F · 2025
Closest in time.
Lmfusion: Adapting pretrained language models for multimodal generation, 2025
Shi, W., Han, X., Zhou, C., Liang, W., Lin, X. V., Zettlemoyer, L., and Yu, L · 2025
Closest in time.