Fetching the paper…
Reading the bibliography…
Multimodal vision language models (VLMs) have made significant progress with the support of continuously increasing model sizes and data volumes.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Clark, C.; Lee, K.; Chang, M.-W.; Kwiatkowski, T.; Collins, M.; and Toutanova, K. 2019 · 1905
Earlier work this paper cites.
Adaptive Mixtures of Local Experts
Jacobs, R. A.; Jordan, M. I.; Nowlan, S. J.; and Hinton, G. E. 1991 · 1991
Earlier work this paper cites.
Improved baselines with momentum contrastive learning. arXiv 2020
Chen, X.; Fan, H.; Girshick, R.; and He, K. 2003 · 2003
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2020 · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Ordonez, V.; Kulkarni, G.; and Berg, T. 2011 · 2011
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Chen, X.; Fang, H.; Lin, T.-Y.; Vedantam, R.; Gupta, S.; Dollár, P.; and Zitnick, C. L. 2015 · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; et al. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
A diagram is worth a dozen images
Kembhavi, A.; Salvato, M.; Kolve, E.; Seo, M.; Hajishirzi, H.; and Farhadi, A. 2016 · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017 · 2017
Earlier work this paper cites.
Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension
Kembhavi, A.; Seo, M.; Schwenk, D.; Choi, J.; Farhadi, A.; and Hajishirzi, H. 2017 · 2017
Earlier work this paper cites.
Dvqa: Understanding data visualizations via question answering
Kafle, K.; Price, B.; Cohen, S.; and Kanan, C. 2018 · 2018
Earlier work this paper cites.
Narayan, S.; Cohen, S. B.; and Lapata, M. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A.; and Manning, C. D. 2019 · 2019
Earlier work this paper cites.
Curved scene text detection via transverse and longitudinal sequence connection
Liu, Y.; Jin, L.; Zhang, S.; Luo, C.; and Zhang, S. 2019 · 2019
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Marino, K.; Rastegari, M.; Farhadi, A.; and Mottaghi, R. 2019 · 2019
Earlier work this paper cites.
Kvqa: Knowledge-aware visual question answering
Shah, S.; Mishra, A.; Yadati, N.; and Talukdar, P. P. 2019 · 2019
Earlier work this paper cites.
Towards vqa models that can read
Singh, A.; Natarajan, V.; Shah, M.; Jiang, Y.; Chen, X.; Batra, D.; Parikh, D.; and Rohrbach, M. 2019 · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y.; Zellers, R.; Gao, J.; Choi, Y.; et al. 2020 · 2020
Earlier work this paper cites.
Unsupervised learning of visual features by contrasting cluster assignments
Caron, M.; Misra, I.; Mairal, J.; Goyal, P.; Bojanowski, P.; and Joulin, A. 2020 · 2020
Earlier work this paper cites.
Movienet: A holistic dataset for movie understanding
Huang, Q.; Xiong, Y.; Rao, A.; Wang, J.; and Lin, D. 2020 · 2020
Earlier work this paper cites.
Textcaps: a dataset for image captioning with reading comprehension
Sidorov, O.; Hu, R.; Rohrbach, M.; and Singh, A. 2020 · 2020
Cited alongside, same era.
Google landmarks dataset v2-a large-scale benchmark for instance-level recognition and retrieval
Weyand, T.; Araujo, A.; Cao, B.; and Sim, J. 2020 · 2020
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
Caron, M.; Touvron, H.; Misra, I.; Jégou, H.; Mairal, J.; Bojanowski, P.; and Joulin, A. 2021 · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Cited alongside, same era.
Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning
Lu, P.; Qiu, L.; Chen, J.; Xia, T.; Zhao, Y.; Zhang, W.; Yu, Z.; Liang, X.; and Zhu, S.-C. 2021 · 2021
Scaling Vision-Language Models with Sparse Mixture of Experts
Shen, S.; Yao, Z.; Li, C.; Darrell, T.; Keutzer, K.; and He, Y. 2023 · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G.; Anil, R.; Borgeaud, S.; Wu, Y.; Alayrac, J.-B.; Yu, J.; Soricut, R.; Schalkwyk, J.; Dai, A. M.; Hauth, A.; et al. 2023 · 2023
Later among the works it cites.
PanGu- π \pi : Enhancing Language Model Architectures via Nonlinearity Compensation
Wang, Y.; Chen, H.; Tang, Y.; Guo, T.; Han, K.; Nie, Y.; Wang, X.; Hu, H.; Bai, Z.; Wang, Y.; et al. 2023 · 2023
Later among the works it cites.
Tinygpt-v: Efficient multimodal large language model via small backbones
Yuan, Z.; Li, Z.; and Sun, L. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Cited alongside, same era.
Imagenet-21k pretraining for the masses
Ridnik, T.; Ben-Baruch, E.; Noy, A.; and Zelnik-Manor, L. 2021 · 2021
Cited alongside, same era.
Scaling Vision with Sparse Mixture of Experts
Riquelme, C.; Puigcerver, J.; Mustafa, B.; Neumann, M.; Jenatton, R.; Pinto, A. S.; Keysers, D.; and Houlsby, N. 2021 · 2021
Cited alongside, same era.
LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Schuhmann, C.; Vencu, R.; Beaumont, R.; Kaczmarczyk, R.; Mullis, C.; Katta, A.; Coombes, T.; Jitsev, J.; and Komatsuzaki, A. 2021 · 2021
Cited alongside, same era.
Document collection visual question answering
Tito, R.; Karatzas, D.; and Valveny, E. 2021 · 2021
Cited alongside, same era.
Fewclue: A chinese few-shot learning evaluation benchmark
Xu, L.; Lu, X.; Yuan, C.; Zhang, X.; Xu, H.; Yuan, H.; Wei, G.; Pan, X.; Tian, X.; Qin, L.; et al. 2021 · 2021
Cited alongside, same era.
VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts
Bao, H.; Wang, W.; Dong, L.; Liu, Q.; Mohammed, O. K.; Aggarwal, K.; Som, S.; and Wei, F. 2022 · 2022
Cited alongside, same era.
Zhai, X.; Mustafa, B.; Kolesnikov, A.; and Beyer, L. 2023 · 2023
Later among the works it cites.
SVIT: Scaling up Visual Instruction Tuning
Zhao, B.; Wu, B.; He, M.; and Huang, T. 2023 · 2023
Later among the works it cites.
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023 · 2023
Later among the works it cites.
Stable lm 2 1.6 b technical report
Bellagente, M.; Tow, J.; Mahan, D.; Phung, D.; Zhuravinskyi, M.; Adithyan, R.; Baicoianu, J.; Brooks, B.; Cooper, N.; Datta, A.; et al. 2024 · 2024
Later among the works it cites.
ALLaVA: Harnessing GPT4V-synthesized Data for A Lite Vision-Language Model
Chen, G. H.; Chen, S.; Zhang, R.; Chen, J.; Wu, X.; Zhang, Z.; Chen, Z.; Li, J.; Wan, X.; and Wang, B. 2024 · 2024
Later among the works it cites.
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Chu, X.; Qiao, L.; Zhang, X.; Xu, S.; Wei, F.; Yang, Y.; Sun, X.; Hu, Y.; Lin, X.; Zhang, B.; and Shen, C. 2024 · 2024
Later among the works it cites.
DeepSeek-Coder: When the Large Language Model Meets Programming–The Rise of Code Intelligence
Guo, D.; Zhu, Q.; Yang, D.; Xie, Z.; Dong, K.; Zhang, W.; Chen, G.; Bi, X.; Wu, Y.; Li, Y.; et al. 2024 · 2024
Later among the works it cites.
C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Huang, Y.; Bai, Y.; Zhu, Z.; Zhang, J.; Zhang, J.; Su, T.; Liu, J.; Lv, C.; Zhang, Y.; Fu, Y.; et al. 2024 · 2024
Later among the works it cites.
Jiang, A. Q.; Sablayrolles, A.; Roux, A.; Mensch, A.; Savary, B.; Bamford, C.; Chaplot, D. S.; Casas, D. d. l.; Hanna, E. B.; Bressand, F.; et al. 2024 · 2024
Later among the works it cites.
Mini-gemini: Mining the potential of multi-modality vision language models
Li, Y.; Zhang, Y.; Wang, C.; Zhong, Z.; Chen, Y.; Chu, R.; Liu, S.; and Jia, J. 2024 · 2024
Later among the works it cites.
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
Lin, B.; Tang, Z.; Ye, Y.; Cui, J.; Zhu, B.; Jin, P.; Huang, J.; Zhang, J.; Ning, M.; and Yuan, L. 2024 · 2024
Later among the works it cites.
Llava-next: Improved reasoning, ocr, and world knowledge
Liu, H.; Li, C.; Li, Y.; Li, B.; Zhang, Y.; Shen, S.; and Lee, Y. J. 2024 · 2024
Later among the works it cites.
DeepSeek-VL: towards real-world vision-language understanding
Lu, H.; Liu, W.; Zhang, B.; Wang, B.; Dong, K.; Liu, B.; Sun, J.; Ren, T.; Li, Z.; Sun, Y.; et al. 2024 · 2024
Later among the works it cites.
Measuring Vision-Language STEM Skills of Neural Models
Shen, J.; Yuan, Y.; Mirzoyan, S.; Zhang, M.; and Wang, C. 2024 · 2024
Later among the works it cites.
Rethinking Optimization and Architecture for Tiny Language Models
Tang, Y.; Liu, F.; Ni, Y.; Tian, Y.; Bai, Z.; Hu, Y.-Q.; Liu, S.; Jui, S.; Han, K.; and Wang, Y. 2024 · 2024
Later among the works it cites.
Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning
Xu, Z.; Feng, C.; Shao, R.; Ashby, T.; Shen, Y.; Jin, D.; Cheng, Y.; Wang, Q.; and Huang, L. 2024 · 2024
Later among the works it cites.
TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones
Yuan, Z.; Li, Z.; Huang, W.; Ye, Y.; and Sun, L. 2024 · 2024
Later among the works it cites.
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Zhang, Y.; Zhang, R.; Gu, J.; Zhou, Y.; Lipka, N.; Yang, D.; and Sun, T. 2024 · 2024
Later among the works it cites.
LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model
Zhu, Y.; Zhu, M.; Liu, N.; Ou, Z.; Mou, X.; and Tang, J. 2024 · 2024
Later among the works it cites.