Fetching the paper…
Reading the bibliography…
Multimodal Large Language Models (MLLMs) have attracted much attention for their multifunctionality.
Signals and systems
A. V. Oppenheim, A. S. Willsky, S. H. Nawab, and J.-J. Ding · 1997
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering, 2017
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people, 2018
D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, et al · 2018
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering, 2019
D. A. Hudson and C. D. Manning · 2019
Earlier work this paper cites.
Decoupled weight decay regularization, 2019
I. Loshchilov and F. Hutter · 2019
Earlier work this paper cites.
Towards vqa models that can read
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach · 2019
Earlier work this paper cites.
Efficiently modeling long sequences with structured state spaces
A. Gu, K. Goel, and C. Ré · 2021
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Earlier work this paper cites.
GLM: general language model pretraining with autoregressive blank infilling
Z. Du, Y. Qian, X. Liu, M. Ding, J. Qiu, Z. Yang, and J. Tang · 2022
Earlier work this paper cites.
It’s raw! audio generation with state-space models, 2022
K. Goel, A. Gu, C. Donahue, and C. Ré · 2022
Earlier work this paper cites.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou · 2023
Earlier work this paper cites.
Shikra: Unleashing multimodal llm’s referential dialogue magic, 2023
K. Chen, Z. Zhang, W. Zeng, R. Zhang, F. Zhu, and R. Zhao · 2023
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing · 2023
Earlier work this paper cites.
Mobilevlm : A fast, strong and open vision language assistant for mobile devices
X. Chu, L. Qiao, X. Lin, S. Xu, Y. Yang, Y. Hu, F. Wei, X. Zhang, B. Zhang, X. Wei, and C. Shen · 2023
Earlier work this paper cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. C. H. Hoi · 2023
Earlier work this paper cites.
Hungry hungry hippos: Towards language modeling with state space models, 2023
D. Y. Fu, T. Dao, K. K. Saab, A. W. Thomas, A. Rudra, and C. Ré · 2023
Earlier work this paper cites.
Llama-adapter v2: Parameter-efficient visual instruction model, 2023
P. Gao, J. Han, R. Zhang, Z. Lin, S. Geng, A. Zhou, et al · 2023
Earlier work this paper cites.
Mamba: Linear-time sequence modeling with selective state spaces, 2023
A. Gu and T. Dao · 2023
Earlier work this paper cites.
Textbooks are all you need, 2023
S. Gunasekar, Y. Zhang, J. Aneja, C. C. T. Mendes, A. D. Giorno, S. Gopi, M. Javaheripi, P. Kauffmann, G. de Rosa, O. Saarikivi, A. Salim, S. Shah, H. S. Behl, X. Wang, S. Bubeck, R. Eldan, A. T. Kalai, Y. T. Lee, and Y. Li · 2023
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
J. Li, D. Li, S. Savarese, and S. C. H. Hoi · 2023
Cited alongside, same era.
Textbooks are all you need ii: phi-1.5
Y. Li, S. Bubeck, R. Eldan, A. Del Giorno, S. Gunasekar, and Y. T. Lee · 2023
Cited alongside, same era.
Evaluating object hallucination in large vision-language models
Y. Li, Y. Du, K. Zhou, J. Wang, W. X. Zhao, and J. rong Wen · 2023
Cited alongside, same era.
Evaluating object hallucination in large vision-language models, 2023
Y. Li, Y. Du, K. Zhou, J. Wang, W. X. Zhao, and J.-R. Wen · 2023
Cited alongside, same era.
Prismatic vlms: Investigating the design space of visually-conditioned language models, 2024
S. Karamcheti, S. Nair, A. Balakrishna, P. Liang, T. Kollar, and D. Sadigh · 2024
Closest in time.
Structure-aware multimodal sequential learning for visual dialog
Y.-J. Kim, M.-J. Kim, K. An, J. Ahn, J. Kim, Y.-J. Heo, D.-S. Chang, and E.-S. Kim · 2024
Closest in time.
Obelics: An open web-scale filtered dataset of interleaved image-text documents
H. Laurençon, L. Saulnier, L. Tronchon, S. Bekman, A. Singh, A. Lozhkov, T. Wang, S. Karamcheti, A. Rush, D. Kiela, et al · 2024
Closest in time.
Mamba-nd: Selective state space modeling for multi-dimensional data
S. Li, H. Singh, and A. Grover · 2024
Closest in time.
Vmamba: Visual state space model
Y. Liu, Y. Tian, Y. Zhao, H. Yu, L. Xie, Y. Wang, Q. Ye, and Y. Liu · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Visual spatial reasoning, 2023
F. Liu, G. Emerson, and N. Collier · 2023
Cited alongside, same era.
Improved baselines with visual instruction tuning, 2023
H. Liu, C. Li, Y. Li, and Y. J. Lee · 2023
Cited alongside, same era.
Visual instruction tuning, 2023
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Cited alongside, same era.
Stanford alpaca: An instruction-following llama model
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models, 2023
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
Cited alongside, same era.
Sigmoid loss for language image pre-training, 2023
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer · 2023
Cited alongside, same era.
Pytorch fsdp: Experiences on scaling fully sharded data parallel, 2023
Y. Zhao, A. Gu, R. Varma, L. Luo, C.-C. Huang, M. Xu, et al · 2023
Cited alongside, same era.
Generative multi-modal knowledge retrieval with large language models
X. Long, J. Zeng, F. Meng, Z. Ma, K. Zhang, B. Zhou, and J. Zhou · 2024
Closest in time.
Gpt-4 technical report, 2024
OpenAI, :, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, et al · 2024
Closest in time.
Dinov2: Learning robust visual features without supervision, 2024
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, et al · 2024
Closest in time.
Efficientvmamba: Atrous selective scan for light weight visual mamba
X. Pei, T. Huang, and C. Xu · 2024
Closest in time.
Vl-mamba: Exploring state space models for multimodal learning, 2024
Y. Qiao, Z. Yu, L. Guo, S. Chen, Z. Zhao, M. Sun, Q. Wu, and J. Liu · 2024
Closest in time.
Autoregressive pretraining with mamba in vision, 2024
S. Ren, X. Li, H. Tu, F. Wang, F. Shu, L. Zhang, J. Mei, L. Yang, P. Wang, H. Wang, A. Yuille, and C. Xie · 2024
Closest in time.
Eyes wide shut? exploring the visual shortcomings of multimodal llms, 2024
S. Tong, Z. Liu, Y. Zhai, Y. Ma, Y. LeCun, and S. Xie · 2024
Closest in time.
T-sciq: Teaching multimodal chain-of-thought reasoning via large language model signals for science question answering
L. Wang, Y. Hu, J. He, X. Xu, N. Liu, H. Liu, and H. T. Shen · 2024
Closest in time.
Altdiffusion: A multilingual text-to-image diffusion model
F. Ye, G. Liu, X. Wu, and L. Wu · 2024
Closest in time.
Tinyllama: An open-source small language model, 2024
P. Zhang, G. Zeng, T. Wang, and W. Lu · 2024
Closest in time.
Cobra: Extending mamba to multi-modal large language model for efficient inference, 2024
H. Zhao, M. Zhang, W. Zhao, P. Ding, S. Huang, and D. Wang · 2024
Closest in time.
Tinyllava: A framework of small-scale large multimodal models, 2024
B. Zhou, Y. Hu, X. Weng, J. Jia, J. Luo, X. Liu, et al · 2024
Closest in time.
Vision mamba: Efficient visual representation learning with bidirectional state space model
L. Zhu, B. Liao, Q. Zhang, X. Wang, W. Liu, and X. Wang · 2024
Closest in time.
Llava-phi: Efficient multi-modal assistant with small language model, 2024
Y. Zhu, M. Zhu, N. Liu, Z. Ou, X. Mou, and J. Tang · 2024
Closest in time.