Fetching the paper…
Reading the bibliography…
We explore Multimodal Large Language Models (MLLMs), which integrate LLMs like GPT-4 to handle multimodal data, including text, images, audio, and more.
D. P. Kingma et al. , “Auto-encoding variational bayes,” arXiv:1312.6114 , 2013
2013
Earlier work this paper cites.
O. Vinyals et al. , “Matching networks for one shot learning,” NeurIPS , vol. 29, 2016
2016
Earlier work this paper cites.
K. He et al. , “Deep residual learning for image recognition,” in CVPR , 2016, pp. 770–778
2016
Earlier work this paper cites.
J. Redmon et al. , “You only look once: Unified, real-time object detection,” in CVPR , 2016, pp. 779–788
2016
Earlier work this paper cites.
A. Van Den Oord et al. , “Neural discrete representation learning,” NeurIPS , vol. 30, 2017
2017
Earlier work this paper cites.
K. He et al. , “Mask r-cnn,” in ICCV , Oct 2017
2017
Earlier work this paper cites.
B. Zhou et al. , “Scene parsing through ade20k dataset,” in CVPR , 2017, pp. 633–641
2017
Earlier work this paper cites.
J. Zhai et al. , “Autoencoder and its various variants,” in SMC . IEEE, 2018, pp. 415–419
2018
Earlier work this paper cites.
J. D. M.-W. C. Kenton et al. , “Bert: Pre-training of deep bidirectional transformers for language understanding,” in NAACL-HLT , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
B. Zagidullin et al. , “Drugcomb: an integrative cancer drug combination data portal,” Nucleic acids research , vol. 47, no. W1, pp. W43–W51, 2019
2019
Earlier work this paper cites.
C. Xu et al. , “Adversarial incomplete multi-view clustering.” in IJCAI , vol. 7, 2019, pp. 3933–3939
2019
Earlier work this paper cites.
Z. Yi et al. , “Incremental learning of gan for detecting multiple adversarial attacks,” in ICANN . Springer, 2019, pp. 673–684
2019
Earlier work this paper cites.
O. Sidorov et al. , “Textcaps: a dataset for image captioning with reading comprehension,” in ECCV . Springer, 2020, pp. 742–758
2020
Earlier work this paper cites.
H. Bao et al. , “Beit: Bert pre-training of image transformers,” in ICLR , 2021
2021
Earlier work this paper cites.
A. Radford et al. , “Learning transferable visual models from natural language supervision,” in ICML , ser. Proceedings of Machine Learning Research, M. Meila and T. Zhang, Eds., vol. 139. PMLR, 18–24 Jul 2021, pp. 8748–8763
2021
Earlier work this paper cites.
A. Dosovitskiy et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR , 2021
2021
Earlier work this paper cites.
A. Radford et al. , “Learning transferable visual models from natural language supervision,” in ICML , 2021
2021
Earlier work this paper cites.
A. Radford et al. , “Learning transferable visual models from natural language supervision,” in ICML , 2021
2021
Earlier work this paper cites.
A. Ramesh et al. , “Zero-shot text-to-image generation,” in ICML . PMLR, 2021, pp. 8821–8831
2021
Earlier work this paper cites.
Y. Yang et al. , “Semi-supervised multi-modal clustering and classification with incomplete modalities,” IEEE Trans. Knowl. Data Eng. , vol. 33, no. 2, pp. 682–695, 2021
2021
Earlier work this paper cites.
W. Zhao, et al. , “Telecomnet: Tag-based weakly-supervised modally cooperative hashing network for image retrieval,” IEEE TPAMI , vol. 44, no. 11, pp. 7940–7954, 2021
2021
Earlier work this paper cites.
X. Du et al. , “Combating word-level adversarial text with robust adversarial training,” in IJCNN . IEEE, 2021, pp. 1–8
2021
Earlier work this paper cites.
Z. Yu et al. , “Dual-encoder transformers with cross-modal alignment for multimodal aspect-based sentiment analysis,” in AACL-IJCNLP . Online only: Association for Computational Linguistics, nov 2022, pp. 414–423
2022
Earlier work this paper cites.
Y. Hao et al. , “Language models are general-purpose interfaces,” arXiv:2206.06336 , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
J.-B. Alayrac et al. , “Flamingo: a visual language model for few-shot learning,” in NeurIPS , S. Koyejo et al. , Eds., vol. 35. Curran Associates, Inc., 2022, pp. 23 716–23 736
2022
Earlier work this paper cites.
D. Lee et al. , “Autoregressive image generation using residual quantization,” in CVPR , 2022, pp. 11 523–11 532
2022
Earlier work this paper cites.
S. Zhang et al. , “Opt: Open pre-trained transformer language models,” arXiv:2205.01068 , 2022
2022
Earlier work this paper cites.
R. Rombach et al. , “High-resolution image synthesis with latent diffusion models,” in CVPR , 2022, pp. 10 684–10 695
2022
Earlier work this paper cites.
H. Liu et al. , “Incorporating multi-source urban data for personalized and context-aware multi-modal transportation recommendation,” IEEE Trans. Knowl. Data Eng. , vol. 34, no. 2, pp. 723–735, 2022
2022
Earlier work this paper cites.
J. Wei et al. , “Emergent abilities of large language models,” arXiv:2206.07682 , Jun 2022
2022
Earlier work this paper cites.
L. Ouyang et al. , “Training language models to follow instructions with human feedback,” NeurIPS , 2022
2022
Earlier work this paper cites.
H. Zhang et al. , “Glipv2: Unifying localization and vision-language understanding,” NeurIPS , vol. 35, pp. 36 067–36 080, 2022
2022
Earlier work this paper cites.
X. Yu et al. , “Point-bert: Pre-training 3d point cloud transformers with masked point modeling,” in CVPR , 2022
2022
Earlier work this paper cites.
P. Hu, “Unsupervised contrastive cross-modal hashing,” IEEE TPAMI , vol. 45, no. 3, pp. 3877–3889, 2022
2022
Earlier work this paper cites.
Z. Han et al. , “Trusted multi-view classification with dynamic evidential fusion,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 45, no. 2, pp. 2551–2566, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
W. Zhao et al. , “A survey of large language models,” arXiv:2303.18223 , 2023
2023
Earlier work this paper cites.
W.-L. Chiang et al. , “Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,” 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
R. Zhang et al. , “Llama-adapter: Efficient fine-tuning of language models with zero-init attention,” Jun 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
B. Peng et al. , “Instruction tuning with gpt-4,” arXiv:2304.03277 , 2023
2023
Earlier work this paper cites.
C. Xu et al. , “Baize: An open-source chat model with parameter-efficient tuning on self-chat data,” ACL , 2023
2023
Earlier work this paper cites.
H. Touvron et al. , “Llama 2: Open foundation and fine-tuned chat models,” arXiv:2307.09288 , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
T. Shen et al. , “Large language model alignment: A survey,” arXiv:2309.15025 , 2023
2023
Earlier work this paper cites.
H. Liu et al. , “Visual instruction tuning,” arXiv prepaint arXiv:2304.08485 , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Y. Koh et al. , “Grounding language models to images for multimodal generation,” ICML , 2023
2023
Earlier work this paper cites.
K. Li et al. , “Videochat: Chat-centric video understanding,” arXiv:2305.06355 , 2023
2023
Earlier work this paper cites.
W. Wang et al. , “Cogvlm: Visual expert for pretrained language models,” NeurIPS , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Zhang et al. , “Next-chat: An lmm for chat, detection and segmentation,” arXiv:2311.04498 , 2023
2023
Earlier work this paper cites.
R. Pi et al. , “Detgpt: Detect what you need via reasoning,” EMNLP , 2023
2023
Earlier work this paper cites.
A. Brohan et al. , “Rt-2: Vision-language-action models transfer web knowledge to robotic control,” PMLR , 2023
2023
Earlier work this paper cites.
X. Zhu et al. , “Pointclip v2: Prompting clip and gpt for powerful 3d open-world learning,” in CVPR , 2023, pp. 2639–2650
2023
Earlier work this paper cites.
R. Zhang et al. , “Prompt, generate, then cache: Cascade of foundation models makes strong few-shot learners,” in CVPR , 2023, pp. 15 211–15 222
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
R. Bavishi et al. , “Introducing our multimodal models,” 2023
2023
Earlier work this paper cites.
Z. Yin et al. , “Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark,” NeurIPS , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Z. Shao, Z. Yu, M. Wang, and J. Yu, “Prompting large language models with answer heuristics for knowledge-based visual question answering,” in Computer Vision and Pattern Recognition (CVPR) , 2023, pp. 14 974–14 983
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
G. Luo et al. , “Cheap and quick: Efficient vision-language instruction tuning for large language models,” NeurIPS , 2023
2023
Earlier work this paper cites.
H. Liu et al. , “Language quantized autoencoders: Towards unsupervised text-image alignment,” NeurIPS , 2023
2023
Earlier work this paper cites.
L. Yu et al. , “Spae: Semantic pyramid autoencoder for multimodal generation with frozen llms,” NeurIPS , 2023
2023
Earlier work this paper cites.
J. Li et al. , “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” ICML , 2023
2023
Earlier work this paper cites.
H. Zhang et al. , “Video-llama: An instruction-tuned audio-visual language model for video understanding,” ACL , 2023
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
S. Yin et al. , “A survey on multimodal large language models,” National Science Review , p. nwae403, 11 2024
2024
Closest in time.
G. Chen et al. , “Videollm: Modeling video sequence with large language models,” CVPR , 2024
2024
Closest in time.
R. Dong et al. , “Dreamllm: Synergistic multimodal comprehension and creation,” ICLR , 2024
2024
Closest in time.
R. Xu et al. , “Pointllm: Empowering large language models to understand point clouds,” ECCV , 2024
2024
Closest in time.
S. Huang et al. , “Language is not all you need: Aligning perception with language models,” NeurIPS , vol. 36, 2024
2024
Closest in time.
W. Hong et al. , “Cogagent: A visual language model for gui agents,” in CVPR , June 2024, pp. 14 281–14 290
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
Y. Huang et al. , “Sparkles: Unlocking chats across multiple images for multimodal instruction-following models,” Aug 2023
2023
Cited alongside, same era.
F. Liu et al. , “Mitigating hallucination in large multi-modal models via robust instruction tuning,” in ICLR , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
W. Wang et al. , “Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,” in NeurIPS , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds. Curran Associates, Inc., 2023, pp. 61 501–61 513
2023
Cited alongside, same era.
Y. Su et al. , “Pandagpt: One model to instruction-follow them all,” ACL , 2023
2023
Cited alongside, same era.
2024
Closest in time.
W. Feng et al. , “Layoutgpt: Compositional visual planning and generation with large language models,” NeurIPS , vol. 36, 2024
2024
Closest in time.
Y. Li et al. , “Stablellava: Enhanced visual instruction tuning with synthesized image-dialogue data,” ACL , 2024
2024
Closest in time.
B.-K. Lee et al. , “Moai: Mixture of all intelligence for large language and vision models,” ECCV , 2024
2024
Closest in time.
W. Hu et al. , “Bliva: A simple multimodal llm for better handling of text-rich visual questions,” AAAI , 2024
2024
Closest in time.
S. Wu et al. , “Next-gpt: Any-to-any multimodal llm,” ICML , 2024
2024
Closest in time.
Y. Zeng et al. , “What matters in training a gpt4-style language model with multimodal inputs?” NAACL , 2024
2024
Closest in time.
H. Rasheed et al. , “Glamm: Pixel grounding large multimodal model,” in CVPR , 2024, pp. 13 009–13 018
2024
Closest in time.
J. Lin et al. , “Vila: On pre-training for visual language models,” in CVPR , 2024, pp. 26 689–26 699
2024
Closest in time.
X. Chen, J. Djolonga et al. , “Pali-x: On scaling up a multilingual vision and language model,” CVPR , 2024
2024
Closest in time.
Z. Chen, J. Wu et al. , “Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,” in CVPR , 2024, pp. 24 185–24 198
2024
Closest in time.
2024
Closest in time.
G. Chen et al. , “Lion: Empowering multimodal large language model with dual-level visual knowledge,” in CVPR , 2024, pp. 26 540–26 550
2024
Closest in time.
X. Lai et al. , “Lisa: Reasoning segmentation via large language model,” in CVPR , 2024, pp. 9579–9589
2024
Closest in time.
2024
Closest in time.
S. Zheng et al. , “Unicode: Learning a unified codebook for multimodal large language models,” 2024
2024
Closest in time.
W. Dai et al. , “Instructblip: Towards general-purpose vision-language models with instruction tuning,” NeurIPS , vol. 36, 2024
2024
Closest in time.
B. He et al. , “Ma-lmm: Memory-augmented large multimodal model for long-term video understanding,” in CVPR , 2024, pp. 13 504–13 514
2024
Closest in time.
X. Zou et al. , “Segment everything everywhere all at once,” in Proceedings of the 37th International Conference on Neural Information Processing Systems , ser. NIPS ’23. Red Hook, NY, USA: Curran Associates Inc., 2024
2024
Closest in time.
Z. Lin et al. , “Sphinx: A mixer of weights, visual embeddings and image scales for multi-modal large language models,” ECCV , 2024
2024
Closest in time.
Q. Ye, H. Xu et al. , “mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,” CVPR , pp. 13 040–13 051, June 2024
2024
Closest in time.
S. Ren et al. , “Timechat: A time-sensitive multimodal large language model for long video understanding,” in CVPR , 2024, pp. 14 313–14 323
2024
Closest in time.
2024
Closest in time.
H. Wei et al. , “Vary: Scaling up the vision vocabulary for large vision-language models,” 2024
2024
Closest in time.
Z. Li et al. , “Monkey: Image resolution and text label are important things for large multi-modal models,” in CVPR , 2024, pp. 26 763–26 773
2024
Closest in time.
Y. Li et al. , “Llama-vid: An image is worth 2 tokens in large language models,” 2024
2024
Closest in time.
K. You et al. , “Ferret-ui: Grounded mobile ui understanding with multimodal llms,” ECCV , 2024
2024
Closest in time.
P. Wu et al. , “V?: Guided visual search as a core mechanism in multimodal llms,” in CVPR , 2024, pp. 13 084–13 094
2024
Closest in time.
R. Gong et al. , “Mindagent: Emergent gaming interaction,” ACL , 2024
2024
Closest in time.
Y. Liang et al. , “Taskmatrix. AI: Completing tasks by connecting foundation models with millions of apis,” Intelligent Computing , vol. 3, p. 0063, 2024
2024
Closest in time.
Z. Gao et al. , “Clova: A closed-loop visual assistant with tool usage and update,” in CVPR , 2024, pp. 13 258–13 268
2024
Closest in time.
2024
Closest in time.
H. Liu et al. , “Improved baselines with visual instruction tuning,” in CVPR , 2024, pp. 26 296–26 306
2024
Closest in time.
Y. Yuan et al. , “Osprey: Pixel understanding with visual instruction tuning,” in CVPR , 2024, pp. 28 202–28 211
2024
Closest in time.
2024
Closest in time.
W. Wang et al. , “The all-seeing project v2: Towards general relation comprehension of the open world,” ECCV , 2024
2024
Closest in time.
H. Rasheed et al. , “Glamm: Pixel grounding large multimodal model,” CVPR , pp. 13 009–13 018, June 2024
2024
Closest in time.
M. Cai et al. , “Vip-llava: Making large multimodal models understand arbitrary visual prompts,” in CVPR , 2024, pp. 12 914–12 923
2024
Closest in time.
H. Guo et al. , “Remote sensing chatgpt: Solving remote sensing tasks with chatgpt and visual models,” IEEE , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
A. Hu et al. , “mplug-paperowl: Scientific diagram analysis with the multimodal large language model,” ACM , 2024
2024
Closest in time.
R. Schumann et al. , “Velma: Verbalization embodiment of llm agents for vision and language navigation in street view,” in AAAI , vol. 38, no. 17, 2024, pp. 18 924–18 933
2024
Closest in time.
J. Han et al. , “Onellm: One framework to align all modalities with language,” in CVPR , June 2024, pp. 26 584–26 595
2024
Closest in time.
J. Zhan et al. , “Anygpt: Unified multimodal llm with discrete sequence modeling,” ACL , 2024
2024
Closest in time.
C. Wu et al. , “Pmc-llama: toward building open-source language models for medicine,” JAMIA , p. ocae045, 2024
2024
Closest in time.
S. Yin et al. , “A survey on multimodal large language models,” IEEE , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. Cha et al. , “Honeybee: Locality-enhanced projector for multimodal llm,” in CVPR , June 2024, pp. 13 817–13 827
2024
Closest in time.
H. W. Chung et al. , “Scaling instruction-finetuned language models,” JMLR , vol. 25, no. 70, pp. 1–53, 2024
2024
Closest in time.
X. Tang et al. , “Medagents: Large language models as collaborators for zero-shot medical reasoning,” ACL , 2024
2024
Closest in time.
T. Tu et al. , “Towards generalist biomedical ai,” NEJM AI , vol. 1, no. 3, p. AIoa2300138, 2024
2024
Closest in time.
2024
Closest in time.
C. Xu et al. , “Reliable conflictive multi-view learning,” AAAI , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Y. Liu et al. , “Mmbench: Is your multi-modal model an all-around player?” ACM , 2024
2024
Closest in time.
K. Haydarov et al. , “Affective visual dialog: A large-scale benchmark for emotional reasoning based on visually grounded conversations,” ACM , 2024
2024
Closest in time.
W.-L. Chiang et al. , “Chatbot arena: An open platform for evaluating llms by human preference,” 2024
2024
Closest in time.
C. Li et al. , “Multimodal foundation models: From specialists to general-purpose assistants,” Foundations and Trends® in Computer Graphics and Vision , vol. 16, no. 1-2, pp. 1–214, 2024
2024
Closest in time.
D. Yang et al. , “Synchronized video storytelling: Generating video narrations with structured storyline,” ACL , 2024
2024
Closest in time.
Y. Kim et al. , “Health-llm: Large language models for health prediction via wearable sensor data,” PMLR , 2024
2024
Closest in time.
S. Cheng et al. , “Can we edit multimodal large language models?” ACL , 2024
2024
Closest in time.
K. Liang, Y. Liu et al. , “Knowledge graph contrastive learning based on relation-symmetrical structure,” IEEE Trans. Knowl. Data Eng. , vol. 36, no. 1, pp. 226–238, 2024
2024
Closest in time.
K. Liang, L. Meng et al. , “A survey of knowledge graph reasoning on graph types: Static, dynamic, and multi-modal,” IEEE TPAMI. , pp. 1–20, 2024
2024
Closest in time.