Fetching the paper…
Reading the bibliography…
In recent years, the application of multimodal large language models (MLLM) in various fields has achieved remarkable success.
GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
Hudson, D. A.; and Manning, C. D. 2019 · 1902
Earlier work this paper cites.
Towards VQA Models That Can Read
Singh, A.; Natarajan, V.; Shah, M.; Jiang, Y.; Chen, X.; et al. 2019 · 1904
Earlier work this paper cites.
Root Mean Square Layer Normalization
Zhang, B.; and Sennrich, R. 2019 · 1910
Earlier work this paper cites.
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
Katharopoulos, A.; Vyas, A.; Pappas, N.; and Fleuret, F. 2020 · 2006
Earlier work this paper cites.
ReferItGame: Referring to Objects in Photographs of Natural Scenes
Kazemzadeh, S.; Ordonez, V.; andre Matten, M.; and Berg, T. L. 2014 · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S.; Agrawal, A.; Lu, J.; Mitchell, M.; Batra, D.; Zitnick, C. L.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; et al. 2016 · 2016
Earlier work this paper cites.
Modeling Context in Referring Expressions
Yu, L.; Poirson, P.; Yang, S.; Berg, A. C.; and Berg, T. L. 2016 · 2016
Earlier work this paper cites.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017 · 2017
Earlier work this paper cites.
VizWiz Grand Challenge: Answering Visual Questions from Blind People
Gurari, D.; Li, Q.; Stangl, A. J.; Guo, A.; Lin, C.; et al. 2018 · 2018
Earlier work this paper cites.
Reflective decoding network for image captioning
Ke, L.; Pei, W.; Li, R.; Shen, X.; and Tai, Y.-W. 2019 · 2019
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2019 · 2019
Earlier work this paper cites.
PyTorch Image Models
Wightman, R. 2019 · 2019
Earlier work this paper cites.
Flamingo: a Visual Language Model for Few-Shot Learning
Alayrac, J.-B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; et al. 2022 · 2022
Earlier work this paper cites.
Efficiently Modeling Long Sequences with Structured State Spaces
Gu, A.; Goel, K.; and Ré, C. 2022 · 2022
Earlier work this paper cites.
Liquid Structural State-Space Models
Hasani, R.; Lechner, M.; Wang, T.-H.; Chahine, M.; Amini, A.; and Rus, D. 2022 · 2022
Earlier work this paper cites.
Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Lu, J.; Clark, C.; Zellers, R.; Mottaghi, R.; and Kembhavi, A. 2022 · 2022
Earlier work this paper cites.
OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Awadalla, A.; Gao, I.; Gardner, J.; Hessel, J.; Hanafy, Y.; et al. 2023 · 2023
Earlier work this paper cites.
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Bai, J.; Bai, S.; Yang, S.; Wang, S.; Tan, S.; et al. 2023 · 2023
Earlier work this paper cites.
Decision S4: Efficient Sequence-Based RL via State Spaces Layers
Bar-David, S.; Zimerman, I.; Nachmani, E.; and Wolf, L. 2023 · 2023
Earlier work this paper cites.
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Brohan, A.; Brown, N.; Carbajal, J.; Chebotar, Y.; Chen, X.; et al. 2023 · 2023
Cited alongside, same era.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; Stoica, I.; and Xing, E. P. 2023 · 2023
Cited alongside, same era.
MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
Chu, X.; Qiao, L.; Lin, X.; Xu, S.; Yang, Y.; Hu, Y.; et al. 2023 · 2023
Cited alongside, same era.
UltraFeedback: Boosting Language Models with High-quality Feedback
Cui, G.; Yuan, L.; Ding, N.; Yao, G.; Zhu, W.; Ni, Y.; et al. 2023 · 2023
Cited alongside, same era.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Wang, J.; Meng, L.; Weng, Z.; He, B.; Wu, Z.; et al. 2023 · 2023
Later among the works it cites.
Diffusion Models Without Attention
Yan, J. N.; Gu, J.; and Rush, A. M. 2023 · 2023
Later among the works it cites.
Sigmoid Loss for Language Image Pre-Training
Zhai, X.; Mustafa, B.; Kolesnikov, A.; and Beyer, L. 2023 · 2023
Later among the works it cites.
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Zhao, Y.; Gu, A.; Varma, R.; Luo, L.; Huang, C.-C.; et al. 2023 · 2023
Later among the works it cites.
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dai, W.; Li, J.; Li, D.; Tiong, A. M. H.; Zhao, J.; et al. 2023 · 2023
Cited alongside, same era.
Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
Ding, N.; Chen, Y.; Xu, B.; Qin, Y.; Zheng, Z.; Hu, S.; et al. 2023 · 2023
Cited alongside, same era.
LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Gao, P.; Han, J.; Zhang, R.; Lin, Z.; Geng, S.; et al. 2023 · 2023
Cited alongside, same era.
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Gu, A.; and Dao, T. 2023 · 2023
Cited alongside, same era.
Scene-centric vs. object-centric image-text cross-modal retrieval: a reproducibility study
Hendriksen, M.; Vakulenko, S.; Kuiper, E.; and de Rijke, M. 2023 · 2023
Cited alongside, same era.
OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents
Laurençon, H.; Saulnier, L.; Tronchon, L.; Bekman, S.; Singh, A.; et al. 2023 · 2023
Cited alongside, same era.
Liu, F.; Emerson, G.; and Collier, N. 2023 · 2023
Cited alongside, same era.
Structured State Space Models for In-Context Reinforcement Learning
Lu, C.; Schroecker, Y.; Gu, A.; Parisotto, E.; Foerster, J.; et al. 2023 · 2023
Cited alongside, same era.
Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023 · 2023
Later among the works it cites.
Stable LM 2 1.6B Technical Report
Bellagente, M.; Tow, J.; Mahan, D.; Phung, D.; Zhuravinskyi, M.; Adithyan, R.; Baicoianu, J.; Brooks, B.; Cooper, N.; Datta, A.; Lee, M.; Mostaque, E.; Pieler, M.; Pinnaparju, N.; Rocha, P.; Saini, H.; Teufel, H.; Zanichelli, N.; and Riquelme, C. 2024 · 2024
Closest in time.
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Chu, X.; Qiao, L.; Zhang, X.; Xu, S.; Wei, F.; Yang, Y.; et al. 2024 · 2024
Closest in time.
Dao, T.; and Gu, A. 2024 · 2024
Closest in time.
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
Ding, P.; Zhao, H.; Song, W.; Zhang, W.; Zhang, M.; et al. 2024 · 2024
Closest in time.
Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models
Karamcheti, S.; Nair, S.; Balakrishna, A.; Liang, P.; Kollar, T.; et al. 2024 · 2024
Closest in time.
OpenVLA: An Open-Source Vision-Language-Action Model
Kim, M. J.; Pertsch, K.; Karamcheti, S.; Xiao, T.; Balakrishna, A.; et al. 2024 · 2024
Closest in time.
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
Lin, B.; Tang, Z.; Ye, Y.; Cui, J.; Zhu, B.; Jin, P.; et al. 2024 · 2024
Closest in time.
Linearizing Large Language Models
Mercat, J.; Vasiljevic, I.; Keh, S.; Arora, K.; Dave, A.; Gaidon, A.; and Kollar, T. 2024 · 2024
Closest in time.
DINOv2: Learning Robust Visual Features without Supervision
Oquab, M.; Darcet, T.; Moutakanni, T.; Vo, H.; Szafraniec, M.; Khalidov, V.; et al. 2024 · 2024
Closest in time.
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
Tong, S.; Liu, Z.; Zhai, Y.; Ma, Y.; LeCun, Y.; and Xie, S. 2024 · 2024
Closest in time.
An Empirical Study of Mamba-based Language Models
Waleffe, R.; Byeon, W.; Riach, D.; Norick, B.; Korthikanti, V.; et al. 2024 · 2024
Closest in time.
NExT-GPT: Any-to-Any Multimodal LLM
Wu, S.; Fei, H.; Qu, L.; Ji, W.; and Chua, T.-S. 2024 · 2024
Closest in time.
TinyLlama: An Open-Source Small Language Model
Zhang, P.; Zeng, G.; Wang, T.; and Lu, W. 2024 · 2024
Closest in time.
TinyLLaVA: A Framework of Small-scale Large Multimodal Models
Zhou, B.; Hu, Y.; Weng, X.; Jia, J.; Luo, J.; Liu, X.; et al. 2024 · 2024
Closest in time.
LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model
Zhu, Y.; Zhu, M.; Liu, N.; Ou, Z.; Mou, X.; and Tang, J. 2024 · 2024
Closest in time.