Fetching the paper…
Reading the bibliography…
The Multi-Modal Large Language Model (MLLM) refers to an extension of the Large Language Model (LLM) equipped with the capability to receive and infer multi-modal data.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Ren, S.; He, K.; Girshick, R.; and Sun, J. 2017 · 2017
Earlier work this paper cites.
Deep learning for generic object detection: A survey
Liu, L.; Ouyang, W.; Wang, X.; Fieguth, P.; Chen, J.; Liu, X.; and Pietikäinen, M. 2020 · 2020
Earlier work this paper cites.
A comprehensive survey of scene graphs: Generation and application
Chang, X.; Ren, P.; Xu, P.; Li, Z.; Chen, X.; and Hauptmann, A. 2021 · 2021
Earlier work this paper cites.
A survey for in-context learning
Dong, Q.; Li, L.; Dai, D.; Zheng, C.; Wu, Z.; Chang, B.; Sun, X.; Xu, J.; and Sui, Z. 2022 · 2022
Earlier work this paper cites.
Panoptic scene graph generation
Yang, J.; Ang, Y. Z.; Guo, Z.; Zhou, K.; Zhang, W.; and Liu, Z. 2022 · 2022
Earlier work this paper cites.
Two-step Masked Language Model for Domain-adapting Multi-modal Task-oriented Dialogue Systems
Chang, Y.; and Ko, Y. 2023 · 2023
Earlier work this paper cites.
Reltr: Relation transformer for scene graph generation
Cong, Y.; Yang, M. Y.; and Rosenhahn, B. 2023 · 2023
Earlier work this paper cites.
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Fu, C.; Chen, P.; Shen, Y.; Qin, Y.; Zhang, M.; Lin, X.; Qiu, Z.; Lin, W.; Yang, J.; Zheng, X.; et al. 2023 · 2023
Earlier work this paper cites.
Llama-adapter v2: Parameter-efficient visual instruction model
Gao, P.; Han, J.; Zhang, R.; Lin, Z.; Geng, S.; Zhou, A.; Zhang, W.; Lu, P.; He, C.; Yue, X.; et al. 2023 · 2023
Earlier work this paper cites.
Multimodal-gpt: A vision and language model for dialogue with humans
Gong, T.; Lyu, C.; Zhang, S.; Wang, Y.; Zheng, M.; Zhao, Q.; Liu, K.; Zhang, W.; Luo, P.; and Chen, K. 2023 · 2023
Cited alongside, same era.
Imagebind-llm: Multi-modality instruction tuning
Han, J.; Zhang, R.; Shao, W.; Gao, P.; Xu, P.; Xiao, H.; Zhang, K.; Liu, C.; Wen, S.; Guo, Z.; et al. 2023 · 2023
Cited alongside, same era.
ChatGPT for shaping the future of dentistry: the potential of multi-modal large language model
Huang, H.; Zheng, O.; Wang, D.; Yin, J.; Wang, Z.; Ding, S.; Yin, H.; Xu, C.; Yang, R.; Zheng, Q.; et al. 2023 · 2023
Cited alongside, same era.
ChatGPT for good? On opportunities and challenges of large language models for education
Kasneci, E.; Seßler, K.; Küchemann, S.; Bannert, M.; Dementieva, D.; Fischer, F.; Gasser, U.; Groh, G.; Günnemann, S.; Hüllermeier, E.; et al. 2023 · 2023
Cited alongside, same era.
Taskmatrix. ai: Completing tasks by connecting foundation models with millions of apis
Large-scale multi-modal pre-trained models: A comprehensive survey
Wang, X.; Chen, G.; Qian, G.; Gao, P.; Wei, X.-Y.; Wang, Y.; Tian, Y.; and Gao, W. 2023 · 2023
Closest in time.
Visual chatgpt: Talking, drawing and editing with visual foundation models
Wu, C.; Yin, S.; Qi, W.; Wang, X.; Tang, Z.; and Duan, N. 2023 · 2023
Closest in time.
mplug-owl: Modularization empowers large language models with multimodality
Ye, Q.; Xu, H.; Xu, G.; Ye, J.; Yan, M.; Zhou, Y.; Wang, J.; Hu, A.; Shi, P.; Shi, Y.; et al. 2023 · 2023
Closest in time.
A Survey on Multimodal Large Language Models
Yin, S.; Fu, C.; Zhao, S.; Li, K.; Sun, X.; Xu, T.; and Chen, E. 2023 · 2023
Closest in time.
Mm-vet: Evaluating large multimodal models for integrated capabilities
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liang, Y.; Wu, C.; Song, T.; Wu, W.; Xia, Y.; Liu, Y.; Ou, Y.; Lu, S.; Ji, L.; Mao, S.; et al. 2023 · 2023
Cited alongside, same era.
Chameleon: Plug-and-play compositional reasoning with large language models
Lu, P.; Peng, B.; Cheng, H.; Galley, M.; Chang, K.-W.; Wu, Y. N.; Zhu, S.-C.; and Gao, J. 2023 · 2023
Cited alongside, same era.
Prompting large language models with answer heuristics for knowledge-based visual question answering
Shao, Z.; Yu, Z.; Wang, M.; and Yu, J. 2023 · 2023
Cited alongside, same era.
Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface
Shen, Y.; Song, K.; Tan, X.; Li, D.; Lu, W.; and Zhuang, Y. 2023 · 2023
Cited alongside, same era.
Pandagpt: One model to instruction-follow them all
Su, Y.; Lan, T.; Li, H.; Xu, J.; Wang, Y.; and Cai, D. 2023 · 2023
Cited alongside, same era.
Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023a
Cited in the paper.
Videochat: Chat-centric video understanding
Li, K.; He, Y.; Wang, Y.; Li, Y.; Wang, W.; Luo, P.; Wang, Y.; Wang, L.; and Qiao, Y. 2023b
Cited in the paper.
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023a
Cited in the paper.
Yu, W.; Yang, Z.; Li, L.; Wang, J.; Lin, K.; Liu, Z.; Wang, X.; and Wang, L. 2023 · 2023
Closest in time.
Enhancing Subtask Performance of Multi-modal Large Language Model
Zhao, Y.; Li, Z.; Zhang, F.; Xu, X.; and Liu, D. 2023 · 2023
Closest in time.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023 · 2023
Closest in time.
Object detection in 20 years: A survey
Zou, Z.; Chen, K.; Shi, Z.; Guo, Y.; and Ye, J. 2023 · 2023
Closest in time.