Fetching the paper…
Reading the bibliography…
Multimodal Large Language Models (MLLMs) have recently shown remarkable perceptual capability in answering visual questions, however, little is known about the limits of their perception.
Pattern matching: The gestalt approach
Ratcliff, J. W., Metzener, D., et al · 1988
Earlier work this paper cites.
VQA: Visual Question Answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D · 2015
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A. and Manning, C. D · 2019
Earlier work this paper cites.
Ocr-vqa: Visual question answering by reading text in images
Mishra, A., Shekhar, S., Singh, A. K., and Chakraborty, A · 2019
Earlier work this paper cites.
Towards vqa models that can read
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Earlier work this paper cites.
Textcaps: a dataset for image captioning with reading comprehension
Sidorov, O., Hu, R., Rohrbach, M., and Singh, A · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Earlier work this paper cites.
LAION-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al · 2022
Earlier work this paper cites.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J · 2023
Cited alongside, same era.
Introducing our multimodal models, 2023
Bavishi, R., Elsen, E., Hawthorne, C., Nye, M., Odena, A., Somani, A., and Taşırlar, S · 2023
Cited alongside, same era.
InstructBLIP: Towards general-purpose vision-language models with instruction tuning, 2023
Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S · 2023
Cited alongside, same era.
Palm-e: An embodied multimodal language model
Driess, D., Xia, F., Sajjadi, M. S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al · 2023
Cited alongside, same era.
LLaMA: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Cogvlm: Visual expert for pretrained language models, 2023
Wang, W., Lv, Q., Yu, W., Hong, W., Qi, J., Wang, Y., Ji, J., Yang, Z., Zhao, L., Song, X., Xu, J., Xu, B., Li, J., Dong, Y., Ding, M., and Tang, J · 2023
Later among the works it cites.
V*: Guided visual search as a core mechanism in multimodal llms
Wu, P. and Xie, S · 2023
Later among the works it cites.
Appagent: Multimodal agents as smartphone users
Yang, Z., Liu, J., Han, Y., Chen, X., Huang, Z., Fu, B., and Yu, G · 2023
Later among the works it cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hong, W., Wang, W., Lv, Q., Xu, J., Yu, W., Ji, J., Wang, Y., Wang, Z., Dong, Y., Ding, M., et al · 2023
Cited alongside, same era.
Large multilingual models pivot zero-shot multimodal learning across languages
Hu, J., Yao, Y., Wang, C., Wang, S., Pan, Y., Chen, Q., Yu, T., Wu, H., Zhao, Y., Zhang, H., Han, X., Lin, Y., Xue, J., Li, D., Liu, Z., and Sun, M · 2023
Cited alongside, same era.
Evaluating object hallucination in large vision-language models
Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, X., and Wen, J.-R · 2023
Cited alongside, same era.
EmbodiedGPT: Vision-language pre-training via embodied chain of thought
Mu, Y., Zhang, Q., Hu, M., Wang, W., Ding, M., Jin, J., Wang, B., Dai, J., Qiao, Y., and Luo, P · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Cited alongside, same era.
BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Li, J., Li, D., Savarese, S., and Hoi, S
Cited in the paper.
Video-llava: Learning united visual representation by alignment before projection
Lin, B., Zhu, B., Ye, Y., Ning, M., Jin, P., and Yuan, L
Cited in the paper.
Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y., et al · 2023
Later among the works it cites.
Visual cropping improves zero-shot question answering of multimodal large language models
Zhang, J., Khayatkhoei, M., Chhikara, P., and Ilievski, F · 2023
Later among the works it cites.
Mmicl: Empowering vision-language model with multi-modal in-context learning
Zhao, H., Cai, Z., Si, S., Ma, X., An, K., Chen, L., Liu, Z., Wang, S., Han, W., and Chang, B · 2023
Later among the works it cites.
The curious case of nonverbal abstract reasoning with multi-modal large language models
Ahrabian, K., Sourati, Z., Sun, K., Zhang, J., Jiang, Y., Morstatter, F., and Pujara, J · 2024
Closest in time.
Eyes wide shut? exploring the visual shortcomings of multimodal llms
Tong, S., Liu, Z., Zhai, Y., Ma, Y., LeCun, Y., and Xie, S · 2024
Closest in time.