Fetching the paper…
Reading the bibliography…
The visual projector, which bridges the vision and language modalities and facilitates cross-modal alignment, serves as a crucial component in MLLMs.
Vizwiz: nearly real-time answers to visual questions
J. P. Bigham, C. Jayant, H. Ji, G. Little, A. Miller, R. C. Miller, R. Miller, A. Tatarowicz, B. White, S. White, et al · 2010
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
V. Ordonez, G. Kulkarni, and T. Berg · 2011
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg · 2014
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio · 2015
Earlier work this paper cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He · 2017
Earlier work this paper cites.
Dvqa: Understanding data visualizations via question answering
K. Kafle, B. Price, S. Cohen, and C. Kanan · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
GQA: A new dataset for real-world visual reasoning and compositional question answering
D. A. Hudson and C. D. Manning · 2019
Earlier work this paper cites.
OK-VQA: A visual question answering benchmark requiring external knowledge
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi · 2019
Earlier work this paper cites.
Ocr-vqa: Visual question answering by reading text in images
A. Mishra, S. Shekhar, A. K. Singh, and A. Chakraborty · 2019
Earlier work this paper cites.
Towards VQA models that can read
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach · 2019
Earlier work this paper cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
E. Voita, D. Talbot, F. Moiseev, R. Sennrich, and I. Titov · 2019
Earlier work this paper cites.
Quantifying attention flow in transformers
S. Abnar and W. Zuidema · 2020
Earlier work this paper cites.
End-to-end object detection with transformers
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Earlier work this paper cites.
Textcaps: a dataset for image captioning with reading comprehension
O. Sidorov, R. Hu, M. Rohrbach, and A. Singh · 2020
Earlier work this paper cites.
Deformable detr: Deformable transformers for end-to-end object detection
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai · 2020
Earlier work this paper cites.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
S. Changpinyo, P. K. Sharma, N. Ding, and R. Soricut · 2021
Earlier work this paper cites.
A survey of convolutional neural networks: analysis, applications, and prospects
Z. Li, F. Liu, W. Yang, S. Peng, and J. Zhou · 2021
Earlier work this paper cites.
Flamingos and hedgehogs in the croquet-ground: Teaching evaluation of NLP systems for undergraduate students
B. Madureira · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Cited alongside, same era.
Learning relation alignment for calibrated cross-modal retrieval
S. Ren, J. Lin, G. Zhao, R. Men, A. Yang, J. Zhou, X. Sun, and H. Yang · 2021
Cited alongside, same era.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
C. Schuhmann, R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki · 2021
Cited alongside, same era.
Vl-interpret: An interactive visualization tool for interpreting vision-language transformers
E. Aflalo, M. Du, S.-Y. Tseng, Y. Liu, C. Wu, N. Duan, and V. Lal · 2022
Cited alongside, same era.
Dinov2: Learning robust visual features without supervision
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al · 2023
Later among the works it cites.
TESTA: Temporal-spatial token aggregation for long-form video-language understanding
S. Ren, S. Chen, S. Li, X. Sun, and L. Hou · 2023
Later among the works it cites.
Delving into the openness of CLIP
S. Ren, L. Li, X. Ren, G. Zhao, and X. Sun · 2023
Later among the works it cites.
Causal interpretation of self-attention in pre-trained transformers
R. Y. Rohekar, Y. Gurwicz, and S. Nisimov · 2023
Later among the works it cites.
Sharegpt
ShareGPT · 2023
Later among the works it cites.
Moviechat: From dense token to sparse memory for long video understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. Binkowski, R. Barreira, O. Vinyals, A. Zisserman, and K. Simonyan · 2022
Cited alongside, same era.
A survey for in-context learning, 2022
Q. Dong, L. Li, D. Dai, C. Zheng, Z. Wu, B. Chang, X. Sun, J. Xu, L. Li, and Z. Sui · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2022
Cited alongside, same era.
Dime: Fine-grained interpretations of multimodal models via disentangled local explanations
Y. Lyu, P. P. Liang, Z. Deng, R. Salakhutdinov, and L.-P. Morency · 2022
Cited alongside, same era.
Investigation of explainability techniques for multimodal transformers
K. Ramesh and Y. S. Koh · 2022
Cited alongside, same era.
A-okvqa: A benchmark for visual question answering using world knowledge
D. Schwenk, A. Khandelwal, C. Clark, K. Marino, and R. Mottaghi · 2022
Cited alongside, same era.
Image difference captioning with pre-training and contrastive learning
L. Yao, W. Wang, and Q. Jin · 2022
Cited alongside, same era.
E. Song, W. Chai, G. Wang, Y. Zhang, H. Zhou, F. Wu, X. Guo, T. Ye, Y. Lu, J.-N. Hwang, et al · 2023
Later among the works it cites.
Edit as you wish: Video description editing with multi-grained commands, 2023
L. Yao, Y. Zhang, Z. Wang, X. Hou, T. Ge, Y. Jiang, and Q. Jin · 2023
Later among the works it cites.
mplug-owl: Modularization empowers large language models with multimodality, 2023
Q. Ye, H. Xu, G. Xu, J. Ye, M. Yan, Y. Zhou, J. Wang, A. Hu, P. Shi, Y. Shi, C. Jiang, C. Li, Y. Xu, H. Chen, J. Tian, Q. Qi, J. Zhang, and F. Huang · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer · 2023
Later among the works it cites.
Mmicl: Empowering vision-language model with multi-modal in-context learning
H. Zhao, Z. Cai, S. Si, X. Ma, K. An, L. Chen, Z. Liu, S. Wang, W. Han, and B. Chang · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny · 2023
Later among the works it cites.
Lvlm-intrepret: An interpretability tool for large vision-language models
G. Ben Melech Stan, R. Yehezkel Rohekar, Y. Gurwicz, M. L. Olson, A. Bhiwandiwalla, E. Aflalo, C. Wu, N. Duan, S.-Y. Tseng, and V. Lal · 2024
Closest in time.
Honeybee: Locality-enhanced projector for multimodal llm
J. Cha, W. Kang, J. Mun, and B. Roh · 2024
Closest in time.
Mllm-bench: Evaluating multimodal llms with per-sample criteria, 2024
W. Ge, S. Chen, G. H. Chen, Z. Chen, J. Chen, S. Yan, C. Zhu, Z. Lin, W. Xie, X. Zhang, Y. Chai, X. Liu, D. Song, X. Wang, A. Gao, Z. Zhang, J. Li, X. Wan, and B. Wang · 2024
Closest in time.
Prismatic vlms: Investigating the design space of visually-conditioned language models
S. Karamcheti, S. Nair, A. Balakrishna, P. Liang, T. Kollar, and D. Sadigh · 2024
Closest in time.
Multimodal arxiv: A dataset for improving scientific comprehension of large vision-language models, 2024
L. Li, Y. Wang, R. Xu, P. Wang, X. Feng, L. Kong, and Q. Liu · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024a
H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee · 2024
Closest in time.
Parameter efficient quasi-orthogonal fine-tuning via givens rotation
X. Ma, X. Chu, Z. Yang, Y. Lin, X. Gao, and J. Zhao · 2024
Closest in time.
Mm1: Methods, analysis & insights from multimodal llm pre-training
B. McKinzie, Z. Gan, J.-P. Fauconnier, S. Dodge, B. Zhang, P. Dufter, D. Shah, X. Du, F. Peng, F. Weers, et al · 2024
Closest in time.
Reka Core, Flash, and Edge: A series of powerful multimodal language models, 2024
T. Reka · 2024
Closest in time.
Prompt pre-training with twenty-thousand classes for open-vocabulary visual recognition
S. Ren, A. Zhang, Y. Zhu, S. Zhang, S. Zheng, M. Li, A. J. Smola, and X. Sun · 2024
Closest in time.
Multimodn—multimodal, multi-task, interpretable modular networks
V. Swamy, M. Satayeva, J. Frej, T. Bossy, T. Vogels, M. Jaggi, T. Käser, and M.-A. Hartley · 2024
Closest in time.