Fetching the paper…
Reading the bibliography…
The rapid evolution of Multimodal Large Language Models (MLLMs) has brought substantial advancements in artificial intelligence, significantly enhancing the capability to understand and generate multimodal content.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in
2009
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in
2014
Earlier work this paper cites.
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen, “Improved techniques for training gans,” in
2016
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma
2017
Earlier work this paper cites.
M. Balog, A. Gaunt, M. Brockschmidt, S. Nowozin, and D. Tarlow, “Deepcoder: Learning to write programs,” in
2017
Earlier work this paper cites.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” in
2017
Earlier work this paper cites.
A. Rohrbach, L. A. Hendricks, K. Burns, T. Darrell, and K. Saenko, “Object hallucination in image captioning,” in
2018
Earlier work this paper cites.
X. Li, X. Yin, C. Li, P. Zhang, X. Hu, L. Zhang, L. Wang, H. Hu, L. Dong, F. Wei
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
OpenAI, “OpenAI: Introducing ChatGPT,” 2022. [Online]. Available:
2022
Earlier work this paper cites.
L. Fan, G. Wang, Y. Jiang, A. Mandlekar, Y. Yang, H. Zhu, A. Tang, D.-A. Huang, Y. Zhu, and A. Anandkumar, “Minedojo: Building open-ended embodied agents with internet-scale knowledge,” in
2022
Earlier work this paper cites.
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds
2022
Earlier work this paper cites.
Y. Li, W. Li, and L. Nie, “Mmcoqa: Conversational question answering over text, tables, and images,” in
2022
Earlier work this paper cites.
S. Lin, J. Hilton, and O. Evans, “Truthfulqa: Measuring how models mimic human falsehoods,” in
2022
Earlier work this paper cites.
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
T. Wang, K. Lin, L. Li, C.-C. Lin, Z. Yang, H. Zhang, Z. Liu, and L. Wang, “Equivariant similarity for vision-language foundation models,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Q. Ye, H. Xu, G. Xu, J. Ye, M. Yan, Y. Zhou, J. Wang, A. Hu, P. Shi, Y. Shi
2023
Earlier work this paper cites.
Y. Liu, H. Duan, Y. Zhang, B. Li, S. Zhang, W. Zhao, Y. Yuan, J. Wang, C. He, Z. Liu
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
W. Zhang, M. Aljunied, C. Gao, Y. K. Chia, and L. Bing, “M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models,” in
2023
Earlier work this paper cites.
J. Van Landeghem, R. Tito, Ł. Borchmann, M. Pietruszka, P. Joziak, R. Powalski, D. Jurkiewicz, M. Coustaty, B. Anckaert, E. Valveny
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Kamath, J. Hessel, and K.-W. Chang, “What’s “up” with vision-language models? investigating their struggle with spatial reasoning,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
S. Cheng, B. Tian, Q. Liu, X. Chen, Y. Wang, H. Chen, and N. Zhang, “Can we edit multimodal large language models?” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Li, K. Pan, Z. Ge, M. Gao, W. Ji, W. Zhang, T.-S. Chua, S. Tang, H. Zhang, and Y. Zhuang, “Fine-tuning multimodal llms to follow zero-shot demonstrative instructions,” in
2023
Earlier work this paper cites.
Y. Bitton, H. Bansal, J. Hessel, R. Shao, W. Zhu, A. Awadalla, J. Gardner, R. Taori, and L. Schimdt, “Visit-bench: a benchmark for vision-language instruction following inspired by real-world use,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
F. Liu, K. Lin, L. Li, J. Wang, Y. Yacoob, and L. Wang, “Mitigating hallucination in large multi-modal models via robust instruction tuning,” in
2023
Earlier work this paper cites.
Z. Sun, S. Shen, S. Cao, H. Liu, C. Li, Y. Shen, C. Gan, L.-Y. Gui, Y.-X. Wang, Y. Yang
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Wang, Y. Zhou, G. Xu, P. Shi, C. Zhao, H. Xu, Q. Ye, M. Yan, J. Zhang, J. Zhu
2023
Earlier work this paper cites.
Y. Li, Y. Du, K. Zhou, J. Wang, X. Zhao, and J.-R. Wen, “Evaluating object hallucination in large vision-language models,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
L. Chen, Y. Zhang, S. Ren, H. Zhao, Z. Cai, Y. Wang, P. Wang, T. Liu, and B. Chang, “Towards end-to-end embodied decision making via multi-modal large language model: Explorations with gpt4-vision and beyond,” 2023
2023
Earlier work this paper cites.
W. Dai, J. Li, D. LI, A. Tiong, J. Zhao, W. Wang, B. Li, P. N. Fung, and S. Hoi, “Instructblip: Towards general-purpose vision-language models with instruction tuning,” in
2023
Earlier work this paper cites.
O. Contributors, “Opencompass: A universal evaluation platform for foundation models,”
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Li, Y. Du, K. Zhou, J. Wang, X. Zhao, and J.-R. Wen, “Evaluating object hallucination in large vision-language models,” in
2023
Earlier work this paper cites.
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P. Fung, “Survey of hallucination in natural language generation,”
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
D. Kelly, Y. Chen, S. E. Cornwell, N. S. Delellis, A. Mayhew, S. Onaolapo, and V. L. Rubin, “Bing chat: the future of search engines?”
2023
Earlier work this paper cites.
J. Betker, G. Goh, L. Jing, T. Brooks, J. Wang, L. Li, L. Ouyang, J. Zhuang, J. Lee, Y. Guo
2023
Earlier work this paper cites.
A. K. M. Abbas, K. Tirumala, D. Simig, S. Ganguli, and A. S. Morcos, “Semdedup: Data-efficient learning at web-scale through semantic deduplication,” in
2023
Earlier work this paper cites.
S. Y. Gadre, G. Ilharco, A. Fang, J. Hayase, G. Smyrnis, T. Nguyen, R. Marten, M. Wortsman, D. Ghosh, J. Zhang
2023
Earlier work this paper cites.
Z. Zhang, H. Wu, E. Zhang, G. Zhai, and W. Lin, “Q-bench: A benchmark for multi-modal foundation models on low-level vision from single images to pairs,”
2024
Earlier work this paper cites.
W. Peng, S. Xie, Z. You, S. Lan, and Z. Wu, “Synthesize diagnose and optimize: Towards fine-grained vision-language understanding,” in
2024
Earlier work this paper cites.
P. Wu and S. Xie, “V: Guided visual search as a core mechanism in multimodal llms,” in
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
S. Tong, Z. Liu, Y. Zhai, Y. Ma, Y. LeCun, and S. Xie, “Eyes wide shut? exploring the visual shortcomings of multimodal llms,” in
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
Z. Liu, F. Fang, X. Feng, X. Du, C. Zhang, Z. Wang, Y. Bai, Q. Zhao, L. Fan, C. Gan
2024
Earlier work this paper cites.
H. P. Zou, V. Samuel, Y. Zhou, W. Zhang, L. Fang, Z. Song, P. S. Yu, and C. Caragea, “Implicitave: An open-source dataset and multimodal llms benchmark for implicit attribute value extraction,” in
2024
Earlier work this paper cites.
Q. Yang, M. Ye, and B. Du, “Emollm: Multimodal emotional understanding meets large language models,”
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
Z. Yin, J. Wang, J. Cao, Z. Shi, D. Liu, M. Li, X. Huang, Z. Wang, L. Sheng, L. Bai
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
B. Li, Y. Ge, Y. Ge, G. Wang, R. Wang, R. Zhang, and Y. Shan, “Seed-bench-2: Benchmarking multimodal large language models,” in
2024
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Y. Zeng, H. Zhang, J. Zheng, J. Xia, G. Wei, Y. Wei, Y. Zhang, T. Kong, and R. Song, “What matters in training a gpt4-style language model with multimodal inputs?” in
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
C. Tran and H. L. Thanh, “Lavy: Vietnamese multimodal large language model,”
2024
Cited alongside, same era.
H.-L. Sun, D.-W. Zhou, Y. Li, S. Lu, C. Yi, Q.-G. Chen, Z. Xu, W. Luo, K. Zhang, D.-C. Zhan
2024
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
H. Hengyuan Zhao, P. Zhou, D. Gao, and M. Z. Shou, “Lova3: Learning to visual question answering, asking and assessment,”
2024
Closest in time.
C. Liu, H. Wu, Y. Zhong, X. Zhang, Y. Wang, and W. Xie, “Intelligent grimm-open-ended visual storytelling via latent diffusion models,” in
2024
Closest in time.
2024
Closest in time.
S. Yun, H. Lin, R. Thushara, M. Q. Bhat, Y. Wang, Z. Jiang, M. Deng, J. Wang, T. Tao, J. Li
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in
2024
Closest in time.
A. Gunjal, J. Yin, and E. Bas, “Detecting and preventing hallucinations in large vision language models,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
S. Cha, J. Lee, Y. Lee, and C. Yang, “Visually dehallucinative instruction generation,” in
2024
Closest in time.
2024
Closest in time.
T. Guan, F. Liu, X. Wu, R. Xian, Z. Li, X. Liu, X. Wang, L. Chen, F. Huang, Y. Yacoob
2024
Closest in time.
2024
Closest in time.
P. Ding, J. Wu, J. Kuang, D. Ma, X. Cao, X. Cai, S. Chen, J. Chen, and S. Huang, “Hallu-pi: Evaluating hallucination in multi-modal large language models within perturbed inputs,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
P. Kaul, Z. Li, H. Yang, Y. Dukler, A. Swaminathan, C. Taylor, and S. Soatto, “Throne: An object-based hallucination benchmark for the free-form generations of large vision-language models,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
M. Zhang and K. Rong, “Automated multi-level preference for mllms,”
2024
Closest in time.
2024
Closest in time.
T. Gu, Z. Zhou, K. Huang, D. Liang, Y. Wang, H. Zhao, Y. Yao, X. Qiao, K. Wang, Y. Yang
2024
Closest in time.
M. Li, L. Li, Y. Yin, M. Ahmed, Z. Liu, and Q. Liu, “Red teaming visual language models,”
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
X. Deng, Y. Gu, B. Zheng, S. Chen, S. Stevens, B. Wang, H. Sun, and Y. Su, “Mind2web: Towards a generalist agent for the web,” in
2024
Closest in time.
C. Rawles, A. Li, D. Rodriguez, O. Riva, and T. Lillicrap, “Androidinthewild: A large-scale dataset for android device control,” in
2024
Closest in time.
J. Y. Koh, R. Lo, L. Jang, V. Duvvur, M. C. Lim, P.-Y. Huang, G. Neubig, S. Zhou, R. Salakhutdinov, and D. Fried, “Visualwebarena: Evaluating multimodal agents on realistic visual web tasks,” in
2024
Closest in time.
T. Xu, L. Chen, D.-J. Wu, Y. Chen, Z. Zhang, X. Yao, Z. Xie, Y. Chen, S. Liu, B. Qian
2024
Closest in time.
2024
Closest in time.
J. Wang, H. Xu, J. Ye, M. Yan, W. Shen, J. Zhang, F. Huang, and J. Sang, “Mobile-agent: Autonomous multi-modal mobile device agent with visual perception,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
A. Majumdar, A. Ajay, X. Zhang, P. Putta, S. Yenamandra, M. Henaff, S. Silwal, P. Mcvay, O. Maksymets, S. Arnaud, K. Yadav, Q. Li, B. Newman, M. Sharma, V. Berges, S. Zhang, P. Agrawal, Y. Bisk, D. Batra, M. Kalakrishnan, F. Meier, C. Paxton, S. Sax, and A. Rajeswaran, “Openeqa: Embodied question answering in the era of foundation models,” in
2024
Closest in time.
2024
Closest in time.
W. Wang, Y. Su, J. Huan, J. Liu, W. Chen, Y. Zhang, C.-Y. Li, K.-J. Chang, X. Xin, L. Shen
2024
Closest in time.
2024
Closest in time.
J. Chen, R. Ouyang, A. Gao, S. Chen, G. H. Chen, X. Wang, R. Zhang, Z. Cai, K. Ji, G. Yu
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
P. Sermanet, T. Ding, J. Zhao, F. Xia, D. Dwibedi, K. Gopalakrishnan, C. Chan, G. Dulac-Arnold, S. Maddineni, N. J. Joshi
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
T. Qian, J. Chen, L. Zhuo, Y. Jiao, and Y.-G. Jiang, “Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.