Fetching the paper…
Reading the bibliography…
Large-scale Vision-Language Models (LVLMs) have significantly advanced with text-aligned vision inputs.
The reviewing of object files: Object-specific integration of information
Kahneman, D.; Treisman, A.; and Gibbs, B. J. 1992 · 1992
Earlier work this paper cites.
Perception and communication
Broadbent, D. E. 2013 · 2013
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Lin, T.-Y.; Maire, M.; Belongie, S.; Bourdev, L.; Girshick, R.; Hays, J.; Perona, P.; Ramanan, D.; Zitnick, C. L.; and Dollár, P. 2015 · 2015
Earlier work this paper cites.
DIML/CVL RGB-D Dataset: 2M RGB-D Images of Natural Indoor and Outdoor Scenes
Cho, J.; Min, D.; Kim, Y.; and Sohn, K. 2021 · 2021
Earlier work this paper cites.
UNIFESP X-ray Body Part Classifier Competition
Eduardo Farina, M. P., FelipeKitamura. 2022 · 2022
Earlier work this paper cites.
Liu, J.; Fan, X.; Huang, Z.; Wu, G.; Liu, R.; Zhong, W.; and Luo, Z. 2022 · 2022
Earlier work this paper cites.
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Bai, J.; Bai, S.; Yang, S.; Wang, S.; Tan, S.; Wang, P.; Lin, J.; Zhou, C.; and Zhou, J. 2023 · 2023
Earlier work this paper cites.
Vision–language model for visual question answering in medical imagery
Bazi, Y.; Rahhal, M. M. A.; Bashmal, L.; and Zuair, M. 2023 · 2023
Earlier work this paper cites.
Imagebind: One embedding space to bind them all
Girdhar, R.; El-Nouby, A.; Liu, Z.; Singh, M.; Alwala, K. V.; Joulin, A.; and Misra, I. 2023 · 2023
Earlier work this paper cites.
Grounding language models to images for multimodal inputs and outputs
Koh, J. Y.; Salakhutdinov, R.; and Fried, D. 2023 · 2023
Earlier work this paper cites.
Gpt-driver: Learning to drive with gpt
Mao, J.; Qian, Y.; Zhao, H.; and Wang, Y. 2023 · 2023
Cited alongside, same era.
thermal dogs and people x6ejw Dataset
Roboflow. 2022 · 2023
Cited alongside, same era.
Pandagpt: One model to instruction-follow them all
Su, Y.; Lan, T.; Li, H.; Xu, J.; Wang, Y.; and Cai, D. 2023 · 2023
Cited alongside, same era.
Cogvlm: Visual expert for pretrained language models
Wang, W.; Lv, Q.; Yu, W.; Hong, W.; Qi, J.; Wang, Y.; Ji, J.; Yang, Z.; Zhao, L.; Song, X.; et al. 2023 · 2023
Cited alongside, same era.
LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models
Xu, P.; Shao, W.; Zhang, K.; Gao, P.; Liu, S.; Lei, M.; Meng, F.; Huang, S.; Qiao, Y.; and Luo, P. 2023 · 2023
Detecting and Preventing Hallucinations in Large Vision Language Models
Gunjal, A.; Yin, J.; and Bas, E. 2024 · 2024
Closest in time.
A Survey on Benchmarks of Multimodal Large Language Models
Li, J.; and Lu, W. 2024 · 2024
Closest in time.
Hello GPT-4o
OpenAI. 2024 · 2024
Closest in time.
InternVL2: Better than the Best—Expanding Performance Boundaries of Open-Source Multimodal Models with the Progressive Scaling Strategy
OpenGVLab. 2024 · 2024
Closest in time.
Timechat: A time-sensitive multimodal large language model for long video understanding
Ren, S.; Yao, L.; Li, S.; Sun, X.; and Hou, L. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
mplug-docowl: Modularized multimodal large language model for document understanding
Ye, J.; Hu, A.; Xu, H.; Ye, Q.; Yan, M.; Dan, Y.; Zhao, C.; Xu, G.; Li, C.; Tian, J.; et al. 2023 · 2023
Cited alongside, same era.
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Yu, W.; Yang, Z.; Li, L.; Wang, J.; Lin, K.; Liu, Z.; Wang, X.; and Wang, L. 2023 · 2023
Cited alongside, same era.
Claude 3.5 sonnet
Anthropic. 2024 · 2024
Cited alongside, same era.
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Chen, Z.; Wang, W.; Tian, H.; Ye, S.; Gao, Z.; Cui, E.; Tong, W.; Hu, K.; Luo, J.; Ma, Z.; et al. 2024 · 2024
Cited alongside, same era.
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Fu, C.; Chen, P.; Shen, Y.; Qin, Y.; Zhang, M.; Lin, X.; Yang, J.; Zheng, X.; Li, K.; Sun, X.; Wu, Y.; and Ji, R. 2024 · 2024
Cited alongside, same era.
TroL: Traversal of Layers for Large Language and Vision Models
Lee, B.-K.; Chung, S.; Kim, C. W.; Park, B.; and Ro, Y. M. 2024a
Cited in the paper.
Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models
Lee, B.-K.; Kim, C. W.; Park, B.; and Ro, Y. M. 2024b
Cited in the paper.
Shi, Y.; Gao, Y.; Lai, Y.; Wang, H.; Feng, J.; He, L.; Wan, J.; Chen, C.; Yu, Z.; and Cao, X. 2024 · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Team, G.; Georgiev, P.; Lei, V. I.; Burnell, R.; Bai, L.; Gulati, A.; Tanzer, G.; Vincent, D.; Pan, Z.; Wang, S.; et al. 2024 · 2024
Closest in time.
Drivegpt4: Interpretable end-to-end autonomous driving via large language model
Xu, Z.; Zhang, Y.; Xie, E.; Zhao, Z.; Guo, Y.; Wong, K.-Y. K.; Li, Z.; and Zhao, H. 2024 · 2024
Closest in time.
What you see is what you read? improving text-image alignment evaluation
Yarom, M.; Bitton, Y.; Changpinyo, S.; Aharoni, R.; Herzig, J.; Lang, O.; Ofek, E.; and Szpektor, I. 2024 · 2024
Closest in time.
Zhang, P.; Dong, X.; Zang, Y.; Cao, Y.; Qian, R.; Chen, L.; Guo, Q.; Duan, H.; Wang, B.; Ouyang, L.; Zhang, S.; Zhang, W.; Li, Y.; Gao, Y.; Sun, P.; Zhang, X.; Li, W.; Li, J.; Wang, W.; Yan, H.; He, C.; Zhang, X.; Chen, K.; Dai, J.; Qiao, Y.; Lin, D.; and Wang, J. 2024 · 2024
Closest in time.