Fetching the paper…
Reading the bibliography…
Large vision-language models (LVLMs) have recently achieved rapid progress, exhibiting great perception and reasoning abilities concerning visual information.
The total variation distance between the binomial and poisson distributions
J. Kennedy and M. Quine · 1989
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele · 2016
Earlier work this paper cites.
Docvqa: A dataset for vqa on document images
M. Mathew, D. Karatzas, and C. Jawahar · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, et al · 2023
Earlier work this paper cites.
Qwen-vl: A frontier large vision-language model with versatile abilities
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou · 2023
Earlier work this paper cites.
Sharegpt4v: Improving large multi-modal models with better captions
L. Chen, J. Li, X. Dong, P. Zhang, C. He, J. Wang, F. Zhao, and D. Lin · 2023
Earlier work this paper cites.
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu, B. Li, P. Luo, T. Lu, Y. Qiao, and J. Dai · 2023
Earlier work this paper cites.
Mme: A comprehensive evaluation benchmark for multimodal large language models
C. Fu, P. Chen, Y. Shen, Y. Qin, M. Zhang, X. Lin, J. Yang, X. Zheng, K. Li, X. Sun, Y. Wu, and R. Ji · 2023
Cited alongside, same era.
Seed-bench: Benchmarking multimodal llms with generative comprehension
B. Li, R. Wang, G. Wang, Y. Ge, Y. Ge, and Y. Shan · 2023
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
J. Li, D. Li, S. Savarese, and S. Hoi · 2023
Cited alongside, same era.
Benchmarking and improving generator-validator consistency of language models
X. L. Li, V. Shrivastava, S. Li, T. Hashimoto, and P. Liang · 2023
Cited alongside, same era.
Generating with confidence: Uncertainty quantification for black-box large language models
Cogvlm: Visual expert for pretrained language models
W. Wang, Q. Lv, W. Yu, W. Hong, J. Qi, Y. Wang, J. Ji, Z. Yang, L. Zhao, X. Song, et al · 2023
Later among the works it cites.
mplug-owl: Modularization empowers large language models with multimodality
Q. Ye, H. Xu, G. Xu, J. Ye, M. Yan, Y. Zhou, J. Wang, A. Hu, P. Shi, Y. Shi, et al · 2023
Later among the works it cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, et al · 2023
Later among the works it cites.
Llama-adapter: Efficient fine-tuning of language models with zero-init attention
R. Zhang, J. Han, C. Liu, P. Gao, A. Zhou, X. Hu, S. Yan, P. Lu, H. Li, and Y. Qiao · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Lin, S. Trivedi, and J. Sun · 2023
Cited alongside, same era.
Improved baselines with visual instruction tuning, 2023
H. Liu, C. Li, Y. Li, and Y. J. Lee · 2023
Cited alongside, same era.
Visual instruction tuning, 2023
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Cited alongside, same era.
Mmbench: Is your multi-modal model an all-around player?
Y. Liu, H. Duan, Y. Zhang, B. Li, S. Zhang, W. Zhao, Y. Yuan, J. Wang, C. He, Z. Liu, et al · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Cited alongside, same era.
X. Bi, D. Chen, G. Chen, S. Chen, D. Dai, C. Deng, H. Ding, K. Dong, Q. Du, Z. Fu, et al · 2024
Closest in time.
Are we on the right way for evaluating large vision-language models?
L. Chen, J. Li, X. Dong, P. Zhang, Y. Zang, Z. Chen, H. Duan, J. Wang, Y. Qiao, D. Lin, et al · 2024
Closest in time.
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Z. Chen, W. Wang, H. Tian, S. Ye, Z. Gao, E. Cui, W. Tong, K. Hu, J. Luo, Z. Ma, et al · 2024
Closest in time.
Mini-gemini: Mining the potential of multi-modality vision language models
Y. Li, Y. Zhang, C. Wang, Z. Zhong, Y. Chen, R. Chu, S. Liu, and J. Jia · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee · 2024
Closest in time.