Fetching the paper…
Reading the bibliography…
Multimodal large language models (MLLMs) have experienced significant advancements recently, but still struggle to recognize and interpret intricate details in high-resolution (HR) images effectively.
Backpropagation applied to handwritten zip code recognition
LeCun, Y.; Boser, B.; Denker, J. S.; Henderson, D.; Howard, R. E.; Hubbard, W.; and Jackel, L. D. 1989 · 1989
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
A Diagram is Worth a Dozen Images
Kembhavi, A.; Salvato, M.; Kolve, E.; Seo, M.; Hajishirzi, H.; and Farhadi, A. 2016 · 2016
Earlier work this paper cites.
Div8k: Diverse 8k resolution image dataset
Gu, S.; Lugmayr, A.; Danelljan, M.; Fritsche, M.; Lamour, J.; and Timofte, R. 2019 · 2019
Earlier work this paper cites.
OCR-VQA: Visual Question Answering by Reading Text in Images
Mishra, A.; Shekhar, S.; Singh, A. K.; and Chakraborty, A. 2019 · 2019
Earlier work this paper cites.
Towards VQA Models That Can Read
Singh, A.; Natarjan, V.; Shah, M.; Jiang, Y.; Chen, X.; Parikh, D.; and Rohrbach, M. 2019 · 2019
Earlier work this paper cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021 · 2021
Earlier work this paper cites.
Docvqa: A dataset for vqa on document images
Mathew, M.; Karatzas, D.; and Jawahar, C. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
GLM: General Language Model Pretraining with Autoregressive Blank Infilling
Du, Z.; Qian, Y.; Liu, X.; Ding, M.; Qiu, J.; Yang, Z.; and Tang, J. 2022 · 2022
Earlier work this paper cites.
Unsupervised Dense Information Retrieval with Contrastive Learning
Izacard, G.; Caron, M.; Hosseini, L.; Riedel, S.; Bojanowski, P.; Joulin, A.; and Grave, E. 2022 · 2022
Earlier work this paper cites.
A convnet for the 2020s
Liu, Z.; Mao, H.; Wu, C.-Y.; Feichtenhofer, C.; Darrell, T.; and Xie, S. 2022 · 2022
Earlier work this paper cites.
Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Lu, P.; Mishra, S.; Xia, T.; Qiu, L.; Chang, K.-W.; Zhu, S.-C.; Tafjord, O.; Clark, P.; and Kalyan, A. 2022 · 2022
Earlier work this paper cites.
ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Masry, A.; Long, D.; Tan, J. Q.; Joty, S.; and Hoque, E. 2022 · 2022
Earlier work this paper cites.
Gpt-4 technical report
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Earlier work this paper cites.
Introducing our Multimodal Models
Bavishi, R.; Elsen, E.; Hawthorne, C.; Nye, M.; Odena, A.; Somani, A.; and Taşırlar, S. 2023 · 2023
Earlier work this paper cites.
Token Merging: Your ViT but Faster
Bolya, D.; Fu, C.-Y.; Dai, X.; Zhang, P.; Feichtenhofer, C.; and Hoffman, J. 2023 · 2023
Cited alongside, same era.
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Fu, C.; Chen, P.; Shen, Y.; Qin, Y.; Zhang, M.; Lin, X.; Yang, J.; Zheng, X.; Li, K.; Sun, X.; et al. 2023 · 2023
Cited alongside, same era.
CORE-MM: Complex Open-Ended Reasoning Evaluation For Multi-Modal Large Language Models
Han, X.; You, Q.; Liu, Y.; Chen, W.; Zheng, H.; Mrini, K.; Lin, X.; Wang, Y.; Zhai, B.; Yuan, J.; Wang, H.; and Yang, H. 2023 · 2023
Cited alongside, same era.
Segment anything
Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A. C.; Lo, W.-Y.; et al. 2023 · 2023
Cited alongside, same era.
OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents
Laurençon, H.; Saulnier, L.; Tronchon, L.; Bekman, S.; Singh, A.; Lozhkov, A.; Wang, T.; Karamcheti, S.; Rush, A.; Kiela, D.; Cord, M.; and Sanh, V. 2023 · 2023
Cited alongside, same era.
VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models
Duan, H.; Yang, J.; Qiao, Y.; Fang, X.; Chen, L.; Liu, Y.; Dong, X.; Zang, Y.; Zhang, P.; Wang, J.; et al. 2024 · 2024
Closest in time.
BLINK: Multimodal Large Language Models Can See but Not Perceive
Fu, X.; Hu, Y.; Li, B.; Feng, Y.; Wang, H.; Lin, X.; Roth, D.; Smith, N. A.; Ma, W.-C.; and Krishna, R. 2024 · 2024
Closest in time.
ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models
Ge, C.; Cheng, S.; Wang, Z.; Yuan, J.; Gao, Y.; Song, J.; Song, S.; Huang, G.; and Zheng, B. 2024 · 2024
Closest in time.
ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
GLM, T.; Zeng, A.; Xu, B.; Wang, B.; Zhang, C.; Yin, D.; Rojas, D.; Feng, G.; Zhao, H.; Lai, H.; Yu, H.; Wang, H.; Sun, J.; Zhang, J.; Cheng, J.; Gui, J.; Tang, J.; Zhang, J.; Li, J.; Zhao, L.; Wu, L.; Zhong, L.; Liu, M.; Huang, M.; Zhang, P.; Zheng, Q.; Lu, R.; Duan, S.; Zhang, S.; Cao, S.; Yang, S.; Tam, W. L.; Zhao, W.; Liu, X.; Xia, X.; Zhang, X.; Gu, X.; Lv, X.; Liu, X.; Liu, X.; Yang, X.; Song, X.; Zhang, X.; An, Y.; Xu, Y.; Niu, Y.; Yang, Y.; Li, Y.; Bai, Y.; Dong, Y.; Qi, Z.; Wang, Z.; Yang, Z.; Du, Z.; Hou, Z.; and Wang, Z. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
CogVLM: Visual Expert for Pretrained Language Models
Wang, W.; Lv, Q.; Yu, W.; Hong, W.; Qi, J.; Wang, Y.; Ji, J.; Yang, Z.; Zhao, L.; Song, X.; Xu, J.; Xu, B.; Li, J.; Dong, Y.; Ding, M.; and Tang, J. 2023 · 2023
Cited alongside, same era.
Vary: Scaling up the vision vocabulary for large vision-language models
Wei, H.; Kong, L.; Chen, J.; Zhao, L.; Ge, Z.; Yang, J.; Sun, J.; Han, C.; and Zhang, X. 2023 · 2023
Cited alongside, same era.
mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video
Xu, H.; Ye, Q.; Yan, M.; Shi, Y.; Ye, J.; Xu, Y.; Li, C.; Bi, B.; Qian, Q.; Wang, W.; Xu, G.; Zhang, J.; Huang, S.; Huang, F.; and Zhou, J. 2023 · 2023
Cited alongside, same era.
Evaluating Object Hallucination in Large Vision-Language Models
Yifan, L.; Yifan, D.; Kun, Z.; Jinpeng, W.; Xin, Z.; and Ji-Rong, W. 2023 · 2023
Cited alongside, same era.
MMBench: Is Your Multi-modal Model an All-around Player?
Yuan, L.; Haodong, D.; Yuanhan, Z.; Bo, L.; Songyang, Z.; Wangbo, Z.; Yike, Y.; Jiaqi, W.; Conghui, H.; Liu, Z.; Kai, C.; and Dahua, L. 2023 · 2023
Cited alongside, same era.
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Zhang, P.; Dong, X.; Wang, B.; Cao, Y.; Xu, C.; Ouyang, L.; Zhao, Z.; Ding, S.; Zhang, S.; Duan, H.; Zhang, W.; Yan, H.; Zhang, X.; Li, W.; Li, J.; Chen, K.; He, C.; Zhang, X.; Qiao, Y.; Lin, D.; and Wang, J. 2023 · 2023
Cited alongside, same era.
Large language models are not robust multiple choice selectors
Zheng, C.; Zhou, H.; Meng, F.; Zhou, J.; and Huang, M. 2023 · 2023
Cited alongside, same era.
Guan, T.; Liu, F.; Wu, X.; Xian, R.; Li, Z.; Liu, X.; Wang, X.; Chen, L.; Huang, F.; Yacoob, Y.; et al. 2024 · 2024
Closest in time.
Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
Luo, G.; Zhou, Y.; Zhang, Y.; Zheng, X.; Sun, X.; and Ji, R. 2024 · 2024
Closest in time.
3AM: An Ambiguity-Aware Multi-Modal Machine Translation Dataset
Ma, X.; Liu, X.; Wong, D. F.; Rao, J.; Li, B.; Ding, L.; Chao, L. S.; Tao, D.; and Zhang, M. 2024 · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M.; Savinov, N.; Teplyashin, D.; Lepikhin, D.; Lillicrap, T.; Alayrac, J.-b.; Soricut, R.; Lazaridou, A.; Firat, O.; Schrittwieser, J.; et al. 2024 · 2024
Closest in time.
Chameleon: Mixed-modal early-fusion foundation models
Team, C. 2024 · 2024
Closest in time.
Eyes wide shut? exploring the visual shortcomings of multimodal llms
Tong, S.; Liu, Z.; Zhai, Y.; Ma, Y.; LeCun, Y.; and Xie, S. 2024 · 2024
Closest in time.
V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs
Wu, P.; and Xie, S. 2024 · 2024
Closest in time.
Yi: Open foundation models by 01. ai
Young, A.; Chen, B.; Li, C.; Huang, C.; Zhang, G.; Zhang, G.; Li, H.; Zhu, J.; Chen, J.; Chang, J.; et al. 2024 · 2024
Closest in time.
Mm-vet: Evaluating large multimodal models for integrated capabilities
Yu, W.; Yang, Z.; Li, L.; Wang, J.; Lin, K.; Liu, Z.; Wang, X.; and Wang, L. 2024 · 2024
Closest in time.
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Yue, X.; Ni, Y.; Zhang, K.; Zheng, T.; Liu, R.; Zhang, G.; Stevens, S.; Jiang, D.; Ren, W.; Sun, Y.; Wei, C.; Yu, B.; Yuan, R.; Sun, R.; Yin, M.; Zheng, B.; Yang, Z.; Liu, Y.; Huang, W.; Sun, H.; Su, Y.; and Chen, W. 2024 · 2024
Closest in time.
Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
Zhang, Y.-F.; Wen, Q.; Fu, C.; Wang, X.; Zhang, Z.; Wang, L.; and Jin, R. 2024 · 2024
Closest in time.