Fetching the paper…
Reading the bibliography…
Recent Multimodal Large Language Models (MLLMs) exhibit impressive abilities to perceive images and follow open-ended instructions.
Microsoft COCO: Common Objects in Context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Visual Dialog
Das, A.; Kottur, S.; Gupta, K.; Singh, A.; Yadav, D.; Moura, J. M.; Parikh, D.; and Batra, D. 2017 · 2017
Earlier work this paper cites.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017 · 2017
Earlier work this paper cites.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; et al. 2017 · 2017
Earlier work this paper cites.
A Minimal Span-Based Neural Constituency Parser
Stern, M.; Andreas, J.; and Klein, D. 2017 · 2017
Earlier work this paper cites.
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-Text Dataset for Automatic Image Captioning
Sharma, P.; Ding, N.; Goodman, S.; and Soricut, R. 2018 · 2018
Earlier work this paper cites.
GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
Hudson, D. A.; and Manning, C. D. 2019 · 2019
Earlier work this paper cites.
OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge
Marino, K.; Rastegari, M.; Farhadi, A.; and Mottaghi, R. 2019 · 2019
Earlier work this paper cites.
Uniter: Universal Image-Text Representation Learning
Chen, Y.-C.; Li, L.; Yu, L.; El Kholy, A.; Ahmed, F.; Gan, Z.; Cheng, Y.; and Liu, J. 2020 · 2020
Earlier work this paper cites.
Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts
Changpinyo, S.; Sharma, P.; Ding, N.; and Soricut, R. 2021 · 2021
Earlier work this paper cites.
Scaling up Visual and Vision-Language Representation Learning with Noisy Text Supervision
Jia, C.; Yang, Y.; Xia, Y.; Chen, Y.-T.; Parekh, Z.; Pham, H.; Le, Q.; Sung, Y.-H.; Li, Z.; and Duerig, T. 2021 · 2021
Earlier work this paper cites.
Align Before Fuse: Vision and Language Representation Learning with Momentum Distillation
Li, J.; Selvaraju, R.; Gotmare, A.; Joty, S.; Xiong, C.; and Hoi, S. C. H. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Florence: A New Foundation Model for Computer Vision
Yuan, L.; Chen, D.; Chen, Y.-L.; Codella, N.; Dai, X.; Gao, J.; Hu, H.; Huang, X.; Li, B.; Li, C.; et al. 2021 · 2021
Cited alongside, same era.
VinVL: Revisiting Visual Representations in Vision-Language Models
Zhang, P.; Li, X.; Hu, X.; Yang, J.; Zhang, L.; Wang, L.; Choi, Y.; and Gao, J. 2021 · 2021
Cited alongside, same era.
Flamingo: a Visual Language Model for Few-Shot Learning
Alayrac, J.-B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; et al. 2022 · 2022
Cited alongside, same era.
CoCa: Contrastive Captioners are Image-Text Foundation Models
Yu, J.; Wang, Z.; Vasudevan, V.; Yeung, L.; Seyedhosseini, M.; and Wu, Y. 2022 · 2022
Later among the works it cites.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; Stoica, I.; and Xing, E. P. 2023 · 2023
Closest in time.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Dai, W.; Li, J.; Li, D.; Tiong, A. M. H.; Zhao, J.; Wang, W.; Li, B.; Fung, P.; and Hoi, S. 2023 · 2023
Closest in time.
PaLM-E: An embodied multimodal language model
Driess, D.; Xia, F.; Sajjadi, M. S.; Lynch, C.; Chowdhery, A.; Ichter, B.; Wahid, A.; Tompson, J.; Vuong, Q.; Yu, T.; et al. 2023 · 2023
Closest in time.
MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts
Bao, H.; Wang, W.; Dong, L.; Liu, Q.; Mohammed, O. K.; Aggarwal, K.; Som, S.; Piao, S.; and Wei, F. 2022 · 2022
Cited alongside, same era.
LAION-COCO: 600m Synthetic Captions from LAION2b-en
Christoph Schuhmann, R. V. T. C. R. B., Andreas Köpf. 2022 · 2022
Cited alongside, same era.
GLM: General Language Model Pretraining with Autoregressive Blank Infilling
Du, Z.; Qian, Y.; Liu, X.; Ding, M.; Qiu, J.; Yang, Z.; and Tang, J. 2022 · 2022
Cited alongside, same era.
A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge
Schwenk, D.; Khandelwal, A.; Clark, C.; Marino, K.; and Mottaghi, R. 2022 · 2022
Cited alongside, same era.
Fengshenbang 1.0: Being the Foundation of Chinese Cognitive Intelligence
Wang, J.; Zhang, Y.; Zhang, L.; Yang, P.; Gao, X.; Wu, Z.; Dong, X.; He, J.; Zhuo, J.; Yang, Q.; Huang, Y.; Li, X.; Wu, Y.; Lu, J.; Zhu, X.; Chen, W.; Han, T.; Pan, K.; Wang, R.; Wang, H.; Wu, X.; Zeng, Z.; Chen, C.; Gan, R.; and Zhang, J. 2022 · 2022
Cited alongside, same era.
Image as a Foreign Language: BEiT Pretraining for Vision and Vision-Language Tasks
Wang, W.; Bao, H.; Dong, L.; Bjorck, J.; Peng, Z.; Liu, Q.; Aggarwal, K.; Mohammed, O. K.; Singhal, S.; Som, S.; et al. 2023a
Cited in the paper.
Self-Instruct: Aligning Language Models with Self-Generated Instructions
Wang, Y.; Kordi, Y.; Mishra, S.; Liu, A.; Smith, N. A.; Khashabi, D.; and Hajishirzi, H. 2023b
Cited in the paper.
Gong, T.; Lyu, C.; Zhang, S.; Wang, Y.; Zheng, M.; Zhao, Q.; Liu, K.; Zhang, W.; Luo, P.; and Chen, K. 2023 · 2023
Closest in time.
Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023 · 2023
Closest in time.
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Ye, Q.; Xu, H.; Xu, G.; Ye, J.; Yan, M.; Zhou, Y.; Wang, J.; Hu, A.; Shi, P.; Shi, Y.; Jiang, C.; Li, C.; Xu, Y.; Chen, H.; Tian, J.; Qi, Q.; Zhang, J.; and Huang, F. 2023 · 2023
Closest in time.
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023 · 2023
Closest in time.