Fetching the paper…
Reading the bibliography…
The advent of Large Language Models (LLMs) has significantly reshaped the trajectory of the AI revolution.
GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
Drew A. Hudson and Christopher D. Manning. 2019 · 1902
Earlier work this paper cites.
Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
Bryan A, Plummer, and et al. 2016 · 2016
Earlier work this paper cites.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Yash Goyal, Tejas Khot, and et al. 2017 · 2017
Earlier work this paper cites.
VizWiz Grand Challenge: Answering Visual Questions from Blind People
Danna, Gurari, and et al. 2018 · 2018
Earlier work this paper cites.
Towards automatic learning of procedures from web instructional videos. In AAAI , Vol. 32
Zhou, Luowei, and et al. 2018 · 2018
Earlier work this paper cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval. In ICCV . 1728–1738
Bain, Max, and et al. 2021 · 2021
Earlier work this paper cites.
Jin, Woojeong, and et al. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision. In ICML . PMLR, 8748–8763
Radford, Alec, and et al. 2021 · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation. In ICML . PMLR, 8821–8831
Ramesh, Aditya, and et al. 2021 · 2021
Earlier work this paper cites.
Multimodal few-shot learning with frozen language models
Tsimpoukelli, Maria, and et al. 2021 · 2021
Earlier work this paper cites.
Simvlm: Simple visual language model pretraining with weak supervision
Wang, Zirui, and et al. 2021 · 2021
Earlier work this paper cites.
Videoclip: Contrastive pre-training for zero-shot video-text understanding
Xu, Hu, and et al. 2021 · 2021
Earlier work this paper cites.
Cm3: A causal masked multimodal model of the internet
Aghajanyan, Armen, and et al. 2022 · 2022
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, Jean-Baptiste, and et al. 2022 · 2022
Earlier work this paper cites.
VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts
Bao, Hangbo, and el al. 2022 · 2022
Earlier work this paper cites.
Pali: A jointly-scaled multilingual language-image model
Chen, Xi, and et al. 2022 · 2022
Earlier work this paper cites.
Gaze estimation using transformer. In ICPR . IEEE, 3341–3347
Cheng, Yihua, and Feng Lu. 2022 · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Lu, Pan, and et al. 2022 · 2022
Earlier work this paper cites.
ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Masry, Ahmed, and et al. 2022 · 2022
Earlier work this paper cites.
Plug-and-play vqa: Zero-shot vqa by conjoining large pretrained models with zero training
Tiong, Anthony Meng Huat, and et al. 2022 · 2022
Earlier work this paper cites.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Tong, Zhan, and et al. 2022 · 2022
Earlier work this paper cites.
Image as a foreign language: Beit pretraining for all vision and vision-language tasks
Wang, Wenhui, and et al. 2022 · 2022
Earlier work this paper cites.
Multiinstruct: Improving multi-modal zero-shot learning via instruction tuning
Xu, Zhiyang, and et al. 2022 · 2022
Earlier work this paper cites.
Video-text modeling with zero-shot transfer from contrastive captioners
Yan, Shen, and et al. 2022 · 2022
Earlier work this paper cites.
An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA
Zhengyuan Yang, Zhe Gan, and et al. 2022 · 2022
Earlier work this paper cites.
Qwen-vl: A frontier large vision-language model with versatile abilities
Bai, Jinze, and et al. 2023b · 2023
Earlier work this paper cites.
Introducing our Multimodal Models
Bavishi, Rohan, and et al. 2023 · 2023
Earlier work this paper cites.
Visit-bench: A benchmark for vision-language instruction following inspired by real-world use
Bitton, Yonatan, and et al. 2023 · 2023
Earlier work this paper cites.
Valor: Vision-audio-language omni-perception pretraining model and dataset
Chen, Sihan, and et al. 2023a · 2023
Earlier work this paper cites.
Chen, Yangyi, and et al. 2023d · 2023
Earlier work this paper cites.
Shikra: Unleashing Multimodal LLM’s Referential Dialogue Magic
Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, and Rui Zhao. 2023e · 2023
Earlier work this paper cites.
Visual Programming for Text-to-Image Generation and Evaluation
Cho, Jaemin, and et al. 2023b · 2023
Earlier work this paper cites.
BharatGPT
corovor.ai. 2023 · 2023
Earlier work this paper cites.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Wenliang Dai and et al. 2023 · 2023
Cited alongside, same era.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning. arXiv 2023
Wenliang Dai, Junnan Li, and et al. 2023 · 2023
Cited alongside, same era.
Palm-e: An embodied multimodal language model
Driess, Danny, and et al. 2023 · 2023
Cited alongside, same era.
PaLM-E: An Embodied Multimodal Language Model
Danny Driess and et al. 2023 · 2023
Cited alongside, same era.
POPE: 6-DoF Promptable Pose Estimation of Any Object, in Any Scene, with One Reference
Knowledge unlearning for llms: Tasks, methods, and challenges
Nianwen Si, Hao Zhang, Heyu Chang, Wenlin Zhang, Dan Qu, and Weiqiang Zhang. 2023 · 2023
Later among the works it cites.
Pandagpt: One model to instruction-follow them all
Su, Yixuan, and et al. 2023 · 2023
Later among the works it cites.
Generative Multimodal Models are In-Context Learners
Sun, Quan, and et al. 2023b · 2023
Later among the works it cites.
Alpha-CLIP: A CLIP Model Focusing on Wherever You Want
Sun, Zeyi, and et al. 2023c · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, Hugo, and et al. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fan, Zhiwen, and et al. 2023 · 2023
Cited alongside, same era.
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Fu, Chaoyou, and et al. 2023 · 2023
Cited alongside, same era.
Imagebind: One embedding space to bind them all. In CVPR . 15180–15190
Girdhar, Rohit, and et al. 2023 · 2023
Cited alongside, same era.
Google Bard
Google. 2023 · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Google, team, and et al. 2023 · 2023
Cited alongside, same era.
From Images to Textual Prompts: Zero-shot Visual Question Answering with Frozen Large Language Models. In CVPR . 10867–10877
Guo, Jiaxian, and et al. 2023 · 2023
Cited alongside, same era.
Bliva: A simple multimodal llm for better handling of text-rich visual questions
Wenbo Hu, Yifan Xu, Yi Li, Weiyue Li, Zeyuan Chen, and Zhuowen Tu. 2023 · 2023
Cited alongside, same era.
Language is not all you need: Aligning perception with language models
Huang, Shaohan, and et al. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
Cogvlm: Visual expert for pretrained language models
Wang, Weihan, and et al. 2023a · 2023
Later among the works it cites.
Large-scale multi-modal pre-trained models: A comprehensive survey
Wang, Xiao, and et al. 2023b · 2023
Later among the works it cites.
Multimodal large language models: A survey
Wu, Jiayang, and et al. 2023b · 2023
Later among the works it cites.
Next-gpt: Any-to-any multimodal llm
Wu, Shengqiong, and et al. 2023c · 2023
Later among the works it cites.
mplug-2: A modularized multi-modal foundation model across text, image and video
Xu, Haiyang, and et al. 2023a · 2023
Later among the works it cites.
Xu, Hu, and et al. 2023b · 2023
Later among the works it cites.
VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners
Shen Yan, Tao Zhu, and et al. 2023 · 2023
Later among the works it cites.
Gpt4tools: Teaching large language model to use tools via self-instruction
Yang, Rui, and et al. 2023a · 2023
Later among the works it cites.
Mm-react: Prompting chatgpt for multimodal reasoning and action
Yang, Zhengyuan, and et al. 2023b · 2023
Later among the works it cites.
mplug-owl: Modularization empowers large language models with multimodality
Ye, Qinghao, and et al. 2023a · 2023
Later among the works it cites.
A Survey on Multimodal Large Language Models
Yin, Shukang, and et al. 2023 · 2023
Later among the works it cites.
Ferret: Refer and ground anything anywhere at any granularity
You, Haoxuan, and et al. 2023 · 2023
Later among the works it cites.
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Yu, Weihao, and et al. 2023 · 2023
Later among the works it cites.
TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones
Yuan, Zhengqing, and et al. 2023b · 2023
Later among the works it cites.
X 2 -VLM: All-In-One Pre-trained Model For Vision-Language Tasks
Zeng, Yan, and et al. 2023 · 2023
Later among the works it cites.
Zhang, Xinsong, and et al. 2023c · 2023
Later among the works it cites.
Chatspot: Bootstrapping multimodal llms via precise referring instruction tuning
Zhao, Liang, and et al. 2023a · 2023
Later among the works it cites.
Chatbridge: Bridging modalities with large language model as a language catalyst
Zhao, Zijia, and et al. 2023c · 2023
Later among the works it cites.
Bubogpt: Enabling visual grounding in multi-modal llms
Yang Zhao, Zhijie Lin, Daquan Zhou, Zilong Huang, Jiashi Feng, and Bingyi Kang. 2023b · 2023
Later among the works it cites.
Minigpt-5: Interleaved vision-and-language generation via generative vokens
Zheng, Kaizhi, and et al. 2023 · 2023
Later among the works it cites.
SkinGPT: A Dermatology Diagnostic System with Vision Large Language Model
Zhou, Juexiao, and et al. 2023 · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, Deyao, and et al. 2023b · 2023
Later among the works it cites.
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
Bin Lin, Zhenyu Tang, Yang Ye, Jiaxi Cui, Bin Zhu, Peng Jin, Junwu Zhang, Munan Ning, and Li Yuan. 2024 · 2024
Closest in time.
PALO: A Polyglot Large Multimodal Model for 5B People
Muhammad Maaz, Hanoona Rasheed, Abdelrahman Shaker, Salman Khan, Hisham Cholakal, Rao M Anwer, Tim Baldwin, Michael Felsberg, and Fahad S Khan. 2024 · 2024
Closest in time.
MoonDream1
vikhyatk. 2024 · 2024
Closest in time.
Mm-llms: Recent advances in multimodal large language models
Duzhen Zhang, Yahan Yu, Chenxing Li, Jiahua Dong, Dan Su, Chenhui Chu, and Dong Yu. 2024 · 2024
Closest in time.
LLaVA-phi: Efficient Multi-Modal Assistant with Small Language Model
Yichen Zhu, Minjie Zhu, Ning Liu, Zhicai Ou, Xiaofeng Mou, and Jian Tang. 2024 · 2024
Closest in time.