Fetching the paper…
Reading the bibliography…
Contrastive Language-Image Pre-training (CLIP) excels in multimodal tasks such as image-text retrieval and zero-shot classification but struggles with fine-grained understanding due to its focus on coarse-grained short captions.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
Kazemzadeh, S., Ordonez, V., Matten, M., and Berg, T · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P., Lai, A., Hodosh, M., and Hockenmaier, J · 2014
Earlier work this paper cites.
Mask r-cnn
He, K., Gkioxari, G., Dollár, P., and Girshick, R · 2017
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Earlier work this paper cites.
Memory attention networks for skeleton-based action recognition
Xie, C., Li, C., Zhang, B., Han, J., Zhen, X., and Chen, J · 2018
Earlier work this paper cites.
Lvis: A dataset for large vocabulary instance segmentation
Gupta, A., Dollar, P., and Girshick, R · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A. and Manning, C. D · 2019
Earlier work this paper cites.
Do imagenet classifiers generalize to imagenet?
Recht, B., Roelofs, R., Schmidt, L., and Shankar, V · 2019
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R · 2020
Earlier work this paper cites.
spacy: Industrial-strength natural language processing in python
Honnibal, M., Montani, I., Van Landeghem, S., Boyd, A., et al · 2020
Earlier work this paper cites.
The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
Kuznetsova, A., Rom, H., Alldrin, N., Uijlings, J., Krasin, I., Pont-Tuset, J., Kamali, S., Popov, S., Malloci, M., Kolesnikov, A., et al · 2020
Earlier work this paper cites.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Changpinyo, S., Sharma, P., Ding, N., and Soricut, R · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A · 2021
Earlier work this paper cites.
Clipcap: Clip prefix for image captioning
Mokady, R., Hertz, A., and Bermano, A. H · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A · 2021
Earlier work this paper cites.
Open-vocabulary object detection using captions
Zareian, A., Rosa, K. D., Hu, D. H., and Chang, S.-F · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Cited alongside, same era.
Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark
Gu, J., Meng, X., Lu, G., Hou, L., Minzhe, N., Liang, X., Yao, L., Huang, R., Zhang, W., Jiang, X., et al · 2022
Cited alongside, same era.
Grounded language-image pre-training
Li, L. H., Zhang, P., Zhang, H., Yang, J., Li, C., Zhong, Y., Wang, L., Yuan, L., Zhang, L., Hwang, J.-N., et al · 2022
Cited alongside, same era.
Learning object-language alignments for open-vocabulary object detection
Lin, C., Sun, P., Jiang, Y., Luo, P., Qu, L., Haffari, G., Yuan, Z., and Cai, J · 2022
Cited alongside, same era.
Cogvlm2: Visual language models for image and video understanding
Hong, W., Wang, W., Ding, M., Yu, W., Lv, Q., Wang, Y., Cheng, Y., Huang, S., Ji, J., Xue, Z., et al · 2024
Later among the works it cites.
Salience detr: Enhancing detection transformer with hierarchical salience filtering refinement
Hou, X., Liu, M., Zhang, S., Wei, P., and Chen, B · 2024
Later among the works it cites.
Fineclip: Self-distilled region-based clip for better fine-grained understanding
Jing, D., He, X., Luo, Y., Fei, N., Yang, G., Wei, W., Zhao, H., and Lu, Z · 2024
Later among the works it cites.
Building and better understanding vision-language models: insights and future directions
Laurençon, H., Marafioti, A., Sanh, V., and Tronchon, L · 2024
Later among the works it cites.
Monkey: Image resolution and text label are important things for large multi-modal models
Li, Z., Yang, B., Liu, Q., Ma, Z., Zhang, S., Yang, J., Sun, Y., Liu, Y., and Bai, X · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Cited alongside, same era.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al · 2022
Cited alongside, same era.
Regionclip: Region-based language-image pretraining
Zhong, Y., Yang, J., Zhang, P., Li, C., Codella, N., Li, L. H., Zhou, L., Dai, X., Yuan, L., Li, Y., et al · 2022
Cited alongside, same era.
Pmc-clip: Contrastive language-image pre-training using biomedical documents
Lin, W., Zhao, Z., Zhang, X., Wu, C., Zhang, Y., Wang, Y., and Xie, W · 2023
Cited alongside, same era.
A prior instruction representation framework for remote sensing image-text retrieval
Pan, J., Ma, Q., and Bai, C · 2023
Cited alongside, same era.
Clip-guided vision-language pre-training for question answering in 3d scenes
Parelli, M., Delitzas, A., Hars, N., Vlassis, G., Anagnostidis, S., Bachmann, G., and Hofmann, T · 2023
Cited alongside, same era.
Eva-clip: Improved training techniques for clip at scale
Sun, Q., Fang, Y., Wu, L., Wang, X., and Cao, Y · 2023
Cited alongside, same era.
Later among the works it cites.
Mmbench: Is your multi-modal model an all-around player?
Liu, Y., Duan, H., Zhang, Y., Li, B., Zhang, S., Zhao, W., Yuan, Y., Wang, J., He, C., Liu, Z., et al · 2024
Later among the works it cites.
Groma: Localized visual tokenization for grounding multimodal large language models
Ma, C., Jiang, Y., Wu, J., Yuan, Z., and Qi, X · 2024
Later among the works it cites.
Scaling open-vocabulary object detection
Minderer, M., Gritsenko, A., and Houlsby, N · 2024
Later among the works it cites.
Kosmos-2: Grounding multimodal large language models to the world
Peng, Z., Wang, W., Dong, L., Hao, Y., Huang, S., Ma, S., Ye, Q., and Wei, F · 2024
Later among the works it cites.
Alpha-clip: A clip model focusing on wherever you want
Sun, Z., Fang, Y., Wu, T., Zhang, P., Zang, Y., Kong, S., Xiong, Y., Lin, D., and Wang, J · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Team, G., Georgiev, P., Lei, V. I., Burnell, R., Bai, L., Gulati, A., Tanzer, G., Vincent, D., Pan, Z., Wang, S., et al · 2024
Later among the works it cites.
A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions
Urbanek, J., Bordes, F., Astolfi, P., Williamson, M., Sharma, V., and Romero-Soriano, A · 2024
Later among the works it cites.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Wang, P., Bai, S., Tan, S., Wang, S., Fan, Z., Bai, J., Chen, K., Liu, X., Wang, J., Ge, W., et al · 2024
Later among the works it cites.
Minicpm-v: A gpt-4v level mllm on your phone
Yao, Y., Yu, T., Zhang, A., Wang, C., Cui, J., Zhu, H., Cai, T., Li, H., Zhao, W., He, Z., et al · 2024
Later among the works it cites.
Long-clip: Unlocking the long-text capability of clip
Zhang, B., Zhang, P., Dong, X., Zang, Y., and Wang, J · 2024
Later among the works it cites.
Dreamlip: Language-image pre-training with long captions
Zheng, K., Zhang, Y., Wu, W., Lu, F., Ma, S., Jin, X., Chen, W., and Shen, Y · 2024
Later among the works it cites.
Iaa: Inner-adaptor architecture empowers frozen large language model with multimodal capabilities
Wang, B., Xie, C., Leng, D., and Yin, Y · 2025
Closest in time.