Fetching the paper…
Reading the bibliography…
Multimodal large language models (MLLMs) have attracted widespread interest and have rich applications.
A new approach to linear filtering and prediction problems
R. E. Kalman · 1960
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2014
Earlier work this paper cites.
Vqa: Visual question answering
A. Agrawal, J. Lu, S. Antol, M. Mitchell, C. L. Zitnick, D. Parikh, and D. Batra · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. M. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Mattnet: Modular attention network for referring expression comprehension
L. Yu, Z. Lin, X. Shen, J. Yang, X. Lu, M. Bansal, and T. L. Berg · 2018
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
D. A. Hudson and C. D. Manning · 2019
Earlier work this paper cites.
Towards vqa models that can read
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach · 2019
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, et al · 2020
Earlier work this paper cites.
Hippo: Recurrent memory with optimal polynomial projections
A. Gu, T. Dao, S. Ermon, A. Rudra, and C. Ré · 2020
Earlier work this paper cites.
Referring expression comprehension: A survey of methods and datasets
Y. Qiao, C. Deng, and Q. Wu · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Earlier work this paper cites.
Combining recurrent, convolutional, and continuous-time models with linear state space layers
A. Gu, I. Johnson, K. Goel, K. Saab, T. Dao, A. Rudra, and C. Ré · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. M. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. C. Hutchinson, R. Pope, J. Bradbury, J. Austin, M. Isard, G. Gur-Ari, P. Yin, T. Duke, A. Levskaya, S. Ghemawat, S. Dev, H. Michalewski, X. García, V. Misra, K. Robinson, L. Fedus, D. Zhou, D. Ippolito, D. Luan, H. Lim, B. Zoph, A. Spiridonov, R. Sepassi, D. Dohan, S. Agrawal, M. Omernick, A. M. Dai, T. S. Pillai, M. Pellat, A. Lewkowycz, E. Moreira, R. Child, O. Polozov, K. Lee, Z. Zhou, X. Wang, B. Saeta, M. Díaz, O. Firat, M. Catasta, J. Wei, K. S. Meier-Hellstern, D. Eck, J. Dean, S. Petrov, and N. Fiedel · 2022
Earlier work this paper cites.
On the parameterization and initialization of diagonal state space models
A. Gu, K. Goel, A. Gupta, and C. Ré · 2022
Earlier work this paper cites.
Efficiently modeling long sequences with structured state spaces
A. Gu, K. Goel, and C. Ré · 2022
Earlier work this paper cites.
Diagonal state spaces are as effective as structured state spaces
A. Gupta, A. Gu, and J. Berant · 2022
Cited alongside, same era.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
J. Li, D. Li, C. Xiong, and S. C. H. Hoi · 2022
Cited alongside, same era.
Learn to explain: Multimodal reasoning via thought chains for science question answering
P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, and A. Kalyan · 2022
Cited alongside, same era.
A-okvqa: A benchmark for visual question answering using world knowledge
D. Schwenk, A. Khandelwal, C. Clark, K. Marino, and R. Mottaghi · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. T. Diab, X. Li, X. V. Lin, T. Mihaylov, M. Ott, S. Shleifer, K. Shuster, D. Simig, P. S. Koura, A. Sridhar, T. Wang, and L. Zettlemoyer · 2022
Evaluating object hallucination in large vision-language models
Y. Li, Y. Du, K. Zhou, J. Wang, W. X. Zhao, and J. rong Wen · 2023
Later among the works it cites.
Improved baselines with visual instruction tuning
H. Liu, C. Li, Y. Li, and Y. J. Lee · 2023
Later among the works it cites.
Visual instruction tuning, 2023
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Later among the works it cites.
Mmbench: Is your multi-modal model an all-around player?
Y. Liu, H. Duan, Y. Zhang, B. Li, S. Zhang, W. Zhao, Y. Yuan, J. Wang, C. He, Z. Liu, K. Chen, and D. Lin · 2023
Later among the works it cites.
Simplified state space layers for sequence modeling
J. T. H. Smith, A. Warrington, and S. W. Linderman · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou · 2023
Cited alongside, same era.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou · 2023
Cited alongside, same era.
Minigpt-v2: large language model as a unified interface for vision-language multi-task learning
J. Chen, D. Zhu, X. Shen, X. Li, Z. Liu, P. Zhang, R. Krishnamoorthi, V. Chandra, Y. Xiong, and M. Elhoseiny · 2023
Cited alongside, same era.
Shikra: Unleashing multimodal llm’s referential dialogue magic
K. Chen, Z. Zhang, W. Zeng, R. Zhang, F. Zhu, and R. Zhao · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing · 2023
Cited alongside, same era.
Mobilevlm : A fast, strong and open vision language assistant for mobile devices
X. Chu, L. Qiao, X. Lin, S. Xu, Y. Yang, Y. Hu, F. Wei, X. Zhang, B. Zhang, X. Wei, and C. Shen · 2023
Cited alongside, same era.
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
Later among the works it cites.
mplug-owl: Modularization empowers large language models with multimodality
Q. Ye, H. Xu, G. Xu, J. Ye, M. Yan, Y. Zhou, J. Wang, A. Hu, P. Shi, Y. Shi, et al · 2023
Later among the works it cites.
Mm-vet: Evaluating large multimodal models for integrated capabilities
W. Yu, Z. Yang, L. Li, J. Wang, K. Lin, Z. Liu, X. Wang, and L. Wang · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny · 2023
Later among the works it cites.
Obelics: An open web-scale filtered dataset of interleaved image-text documents
H. Laurençon, L. Saulnier, L. Tronchon, S. Bekman, A. Singh, A. Lozhkov, T. Wang, S. Karamcheti, A. Rush, D. Kiela, et al · 2024
Closest in time.
Vmamba: Visual state space model
Y. Liu, Y. Tian, Y. Zhao, H. Yu, L. Xie, Y. Wang, Q. Ye, and Y. Liu · 2024
Closest in time.
U-mamba: Enhancing long-range dependency for biomedical image segmentation
J. Ma, F. Li, and B. Wang · 2024
Closest in time.
Vm-unet: Vision mamba unet for medical image segmentation
J. Ruan and S. Xiang · 2024
Closest in time.
Segmamba: Long-range sequential modeling mamba for 3d medical image segmentation
Z.-Y. Xing, T. Ye, Y. Yang, G. Liu, and L. Zhu · 2024
Closest in time.
Vivim: a video vision mamba for medical video object segmentation
Y. Yang, Z.-Y. Xing, and L. Zhu · 2024
Closest in time.
Vision mamba: Efficient visual representation learning with bidirectional state space model
L. Zhu, B. Liao, Q. Zhang, X. Wang, W. Liu, and X. Wang · 2024
Closest in time.
Llava-phi: Efficient multi-modal assistant with small language model
Y. Zhu, M. Zhu, N. Liu, Z. Ou, X. Mou, and J. Tang · 2024
Closest in time.