Fetching the paper…
Reading the bibliography…
Multimodal Large Language Models (MLLMs) have experienced significant advancements recently.
Vqa: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server, 2015
Chen, X., Fang, H., Lin, T.-Y., Vedantam, R., Gupta, S., Dollar, P., and Zitnick, C. L · 2015
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.-J., Shamma, D. A., et al · 2017
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
Gurari, D., Li, Q., Stangl, A. J., Guo, A., Lin, C., Grauman, K., Luo, J., and Bigham, J. P · 2018
Earlier work this paper cites.
Scene text visual question answering
Biten, A. F., Tito, R., Mafla, A., Gomez, L., Rusinol, M., Valveny, E., Jawahar, C., and Karatzas, D · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A. and Manning, C. D · 2019
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Marino, K., Rastegari, M., Farhadi, A., and Mottaghi, R · 2019
Earlier work this paper cites.
Ocr-vqa: Visual question answering by reading text in images
Mishra, A., Shekhar, S., Singh, A. K., and Chakraborty, A · 2019
Earlier work this paper cites.
Towards vqa models that can read
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M · 2019
Earlier work this paper cites.
Captioning images taken by people who are blind
Gurari, D., Zhao, Y., Zhang, M., and Bhattacharya, N · 2020
Earlier work this paper cites.
Textcaps: a dataset for image captioning with reading comprehension
Sidorov, O., Hu, R., Rohrbach, M., and Singh, A · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Earlier work this paper cites.
Masked autoencoders are scalable vision learners, 2021
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2021
Earlier work this paper cites.
Docvqa: A dataset for vqa on document images
Mathew, M., Karatzas, D., and Jawahar, C · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Earlier work this paper cites.
Translating math formula images to latex sequences using deep neural networks with sequence-level training
Wang, Z. and Liu, J.-C · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Cited alongside, same era.
Deepspeed inference: Enabling efficient inference of transformer models at unprecedented scale, 2022
Aminabadi, R. Y., Rajbhandari, S., Zhang, M., Awan, A. A., Li, C., Li, D., Zheng, E., Rasley, J., Smith, S., Ruwase, O., and He, Y · 2022
Cited alongside, same era.
Coyo-700m: Image-text pair dataset
Byeon, M., Park, B., Kim, H., Lee, S., Baek, W., and Kim, S · 2022
Cited alongside, same era.
Unified pretraining framework for document understanding, 2022
Gu, J., Kuen, J., Morariu, V. I., Zhao, H., Barmpalios, N., Jain, R., Nenkova, A., and Sun, T · 2022
Cited alongside, same era.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Lu, P., Mishra, S., Xia, T., Qiu, L., Chang, K.-W., Zhu, S.-C., Tafjord, O., Clark, P., and Kalyan, A · 2022
Cited alongside, same era.
Infimm-eval: Complex open-ended reasoning evaluation for multi-modal large language models, 2023
Han, X., You, Q., Liu, Y., Chen, W., Zheng, H., Mrini, K., Lin, X., Wang, Y., Zhai, B., Yuan, J., Wang, H., and Yang, H · 2023
Later among the works it cites.
From clip to dino: Visual encoders shout in multi-modal large language models, 2023
Jiang, D., Liu, Y., Liu, S., Zhang, X., Li, J., Xiong, H., and Tian, Q · 2023
Later among the works it cites.
Obelics: An open web-scale filtered dataset of interleaved image-text documents, 2023
Laurençon, H., Saulnier, L., Tronchon, L., Bekman, S., Singh, A., Lozhkov, A., Wang, T., Karamcheti, S., Rush, A. M., Kiela, D., Cord, M., and Sanh, V · 2023
Later among the works it cites.
Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models, 2023
Lin, Z., Liu, C., Zhang, R., Gao, P., Qiu, L., Xiao, H., Qiu, H., Lin, C., Shao, W., Chen, K., Han, J., Huang, S., Zhang, Y., He, X., Li, H., and Qiao, Y · 2023
Later among the works it cites.
Eva-clip: Improved training techniques for clip at scale, 2023
Sun, Q., Fang, Y., Wu, L., Wang, X., and Cao, Y · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A-okvqa: A benchmark for visual question answering using world knowledge
Schwenk, D., Khandelwal, A., Clark, C., Marino, K., and Mottaghi, R · 2022
Cited alongside, same era.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023
Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J · 2023
Cited alongside, same era.
Introducing our multimodal models, 2023
Bavishi, R., Elsen, E., Hawthorne, C., Nye, M., Odena, A., Somani, A., and Taşırlar, S · 2023
Cited alongside, same era.
Honeybee: Locality-enhanced projector for multimodal llm, 2023
Cha, J., Kang, W., Mun, J., and Roh, B · 2023
Cited alongside, same era.
Shikra: Unleashing multimodal llm’s referential dialogue magic, 2023
Chen, K., Zhang, Z., Zeng, W., Zhang, R., Zhu, F., and Zhao, R · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Cited alongside, same era.
Palm-e: An embodied multimodal language model, 2023
Driess, D., Xia, F., Sajjadi, M. S. M., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., Huang, W., Chebotar, Y., Sermanet, P., Duckworth, D., Levine, S., Vanhoucke, V., Hausman, K., Toussaint, M., Greff, K., Zeng, A., Mordatch, I., and Florence, P · 2023
Cited alongside, same era.
Later among the works it cites.
Cogvlm: Visual expert for pretrained language models, 2023
Wang, W., Lv, Q., Yu, W., Hong, W., Qi, J., Wang, Y., Ji, J., Yang, Z., Zhao, L., Song, X., Xu, J., Xu, B., Li, J., Dong, Y., Ding, M., and Tang, J · 2023
Later among the works it cites.
The dawn of lmms: Preliminary explorations with gpt-4v(ision), 2023
Yang, Z., Li, L., Lin, K., Wang, J., Lin, C.-C., Liu, Z., and Wang, L · 2023
Later among the works it cites.
mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration, 2023
Ye, Q., Xu, H., Ye, J., Yan, M., Hu, A., Liu, H., Qian, Q., Zhang, J., Huang, F., and Zhou, J · 2023
Later among the works it cites.
Mm-vet: Evaluating large multimodal models for integrated capabilities, 2023
Yu, W., Yang, Z., Li, L., Wang, J., Lin, K., Liu, Z., Wang, X., and Wang, L · 2023
Later among the works it cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi, 2023
Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y., Wei, C., Yu, B., Yuan, R., Sun, R., Yin, M., Zheng, B., Yang, Z., Liu, Y., Huang, W., Sun, H., Su, Y., and Chen, W · 2023
Later among the works it cites.
Halle-switch: Controlling object hallucination in large vision language models, 2023
Zhai, B., Yang, S., Xu, C., Shen, S., Keutzer, K., and Li, M · 2023
Later among the works it cites.
Llavar: Enhanced visual instruction tuning for text-rich image understanding, 2023
Zhang, Y., Zhang, R., Gu, J., Zhou, Y., Lipka, N., Yang, D., and Sun, T · 2023
Later among the works it cites.
Coco is ”all” you need for visual instruction fine-tuning, 2024
Han, X., Wang, Y., Zhai, B., You, Q., and Yang, H · 2024
Closest in time.
Eyes wide shut? exploring the visual shortcomings of multimodal llms, 2024
Tong, S., Liu, Z., Zhai, Y., Ma, Y., LeCun, Y., and Xie, S · 2024
Closest in time.
Exploring the reasoning abilities of multimodal large language models (mllms): A comprehensive survey on emerging trends in multimodal reasoning, 2024
Wang, Y., Chen, W., Han, X., Lin, X., Zhao, H., Liu, Y., Zhai, B., Yuan, J., You, Q., and Yang, H · 2024
Closest in time.