Fetching the paper…
Reading the bibliography…
After pre-training on extensive image-text pairs, Contrastive Language-Image Pre-training (CLIP) demonstrates promising performance on a wide variety of benchmarks.
Efficiency of a good but not linear set union algorithm
Robert Endre Tarjan. 1975 · 1975
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation. In ACL
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Latent dirichlet allocation
David M Blei, Andrew Y Ng, and Michael I Jordan. 2003 · 2003
Earlier work this paper cites.
NLTK: The Natural Language Toolkit. In ACL
Steven Bird and Edward Loper. 2004 · 2004
Earlier work this paper cites.
Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories. In CVPR
Li Fei-Fei, Rob Fergus, and Pietro Perona. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In ACL
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Automated flower classification over a large number of classes. In Sixth Indian Conference on Computer Vision, Graphics & Image Processing
Maria-Elena Nilsback and Andrew Zisserman. 2008 · 2008
Earlier work this paper cites.
Visualizing data using t-SNE
Laurens Van der Maaten and Geoffrey Hinton. 2008 · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database. In CVPR
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo. In ICCV
Jianxiong Xiao, James Hays, Krista A Ehinger, Aude Oliva, and Antonio Torralba. 2010 · 2010
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning. In AISTATES
Adam Coates, Andrew Ng, and Honglak Lee. 2011 · 2011
Earlier work this paper cites.
Are we ready for autonomous driving? the kitti vision benchmark suite. In CVPR
Andreas Geiger, Philip Lenz, and Raquel Urtasun. 2012 · 2012
Earlier work this paper cites.
Cats and dogs. In ICCV
Omkar M Parkhi, Andrea Vedaldi, Andrew Zisserman, and CV Jawahar. 2012 · 2012
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. 2012 · 2012
Earlier work this paper cites.
3d object representations for fine-grained categorization. In ICCVW
Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 2013 · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft
Subhransu Maji, Esa Rahtu, Juho Kannala, Matthew Blaschko, and Andrea Vedaldi. 2013 · 2013
Earlier work this paper cites.
Birdsnap: Large-scale fine-grained visual categorization of birds. In CVPR
Thomas Berg, Jiongxin Liu, Seung Woo Lee, Michelle L Alexander, David W Jacobs, and Peter N Belhumeur. 2014 · 2014
Earlier work this paper cites.
Food-101–mining discriminative components with random forests. In ECCV
Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. 2014 · 2014
Earlier work this paper cites.
Describing textures in the wild. In CVPR
Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context. In ECCV
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick. 2015 · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation. In CVPR
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition. In CVPR
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
YFCC100M: The new data in multimedia research
Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li. 2016 · 2016
Earlier work this paper cites.
Remote sensing image scene classification: Benchmark and state of the art
Gong Cheng, Junwei Han, and Xiaoqiang Lu. 2017 · 2017
Earlier work this paper cites.
Simple and effective multi-paragraph reading comprehension. In ACL
Christopher Clark and Matt Gardner. 2017 · 2017
Cited alongside, same era.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In CVPR
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017 · 2017
Cited alongside, same era.
Multimodal machine learning: A survey and taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency. 2018 · 2018
Cited alongside, same era.
Deep multimodal representation learning: A survey
Wenzhong Guo, Jianwen Wang, and Shiping Wang. 2019 · 2019
Cited alongside, same era.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth. 2019 · 2019
Cited alongside, same era.
Billion-scale similarity search with GPUs
Gridclip: One-stage object detection by grid-level clip representation learning
Jiayi Lin and Shaogang Gong. 2023 · 2023
Later among the works it cites.
Clip is also an efficient segmenter: A text-driven approach for weakly supervised semantic segmentation. In CVPR . 15305–15314
Yuqi Lin, Minghao Chen, Wenxiao Wang, Boxi Wu, Ke Li, Binbin Lin, Haifeng Liu, and Xiaofei He. 2023 · 2023
Later among the works it cites.
Filtering, distillation, and hard negatives for vision-language pre-training. In CVPR
Filip Radenovic, Abhimanyu Dubey, Abhishek Kadian, Todor Mihaylov, Simon Vandenhende, Yash Patel, Yi Wen, Vignesh Ramanathan, and Dhruv Mahajan. 2023 · 2023
Later among the works it cites.
Eva-clip: Improved training techniques for clip at scale
Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. 2023 · 2023
Later among the works it cites.
Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching. In CVPR . 7031–7040
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019 · 2019
Cited alongside, same era.
Decoupled Weight Decay Regularization. In ICLR
I Loshchilov. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
A simple framework for contrastive learning of visual representations. In ICML
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020 · 2020
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2020
Cited alongside, same era.
The hateful memes challenge: Detecting hate speech in multimodal memes. In NeurIPS
Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ringshia, and Davide Testuggine. 2020 · 2020
Cited alongside, same era.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In CVPR
Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. 2021 · 2021
Cited alongside, same era.
Xiaoshi Wu, Feng Zhu, Rui Zhao, and Hongsheng Li. 2023 · 2023
Later among the works it cites.
Alip: Adaptive language-image pre-training with synthetic caption. In ICCV
Kaicheng Yang, Jiankang Deng, Xiang An, Jiawei Li, Ziyong Feng, Jia Guo, Jing Yang, and Tongliang Liu. 2023 · 2023
Later among the works it cites.
Zegclip: Towards adapting clip for zero-shot semantic segmentation. In CVPR . 11175–11185
Ziqin Zhou, Yinjie Lei, Bowen Zhang, Lingqiao Liu, and Yifan Liu. 2023 · 2023
Later among the works it cites.
ShareGPT4V: Improving Large Multi-Modal Models with Better Captions. In ECCV
Lin Chen, Jisong Li, Xiaoyi Dong, Pan Zhang, Conghui He, Jiaqi Wang, Feng Zhao, and Dahua Lin. 2024 · 2024
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Later among the works it cites.
Improving clip training with language rewrites. In NeurIPS
Lijie Fan, Dilip Krishnan, Phillip Isola, Dina Katabi, and Yonglong Tian. 2024 · 2024
Later among the works it cites.
Datacomp: In search of the next generation of multimodal datasets. In NeurIPS
Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, et al · 2024
Later among the works it cites.
Rwkv-clip: A robust vision-language representation learner. In EMNLP
Tiancheng Gu, Kaicheng Yang, Xiang An, Ziyong Feng, Dongnan Liu, Weidong Cai, and Jiankang Deng. 2024 · 2024
Later among the works it cites.
VeCLIP: Improving CLIP Training via Visual-enriched Captions. In ECCV
Zhengfeng Lai, Haotian Zhang, Bowen Zhang, Wentao Wu, Haoping Bai, Aleksei Timofeev, Xianzhi Du, Zhe Gan, Jiulong Shan, Chen-Nee Chuah, Yinfei Yang, and Meng Cao. 2024 · 2024
Later among the works it cites.
Obelics: An open web-scale filtered dataset of interleaved image-text documents. In NeurIPS
Hugo Laurençon, Lucile Saulnier, Léo Tronchon, Stas Bekman, Amanpreet Singh, Anton Lozhkov, Thomas Wang, Siddharth Karamcheti, Alexander Rush, Douwe Kiela, et al · 2024
Later among the works it cites.
Scaling (down) clip: A comprehensive analysis of data, architecture, and training strategies
Zichao Li, Cihang Xie, and Ekin Dogus Cubuk. 2024a · 2024
Later among the works it cites.
Improved baselines with visual instruction tuning. In CVPR
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2024 · 2024
Later among the works it cites.
Transductive Zero-Shot and Few-Shot CLIP. In CVPR
Ségolène Martin, Yunshi Huang, Fereshteh Shakeri, Jean-Christophe Pesquet, and Ismail Ben Ayed. 2024 · 2024
Later among the works it cites.
DeIL: Direct-and-Inverse CLIP for Open-World Few-Shot Learning. In CVPR
Shuai Shao, Yu Bai, Yan Wang, Baodi Liu, and Yicong Zhou. 2024 · 2024
Later among the works it cites.
AMU-Tuning: Effective Logit Bias for CLIP-based Few-shot Learning. In CVPR
Yuwei Tang, Zhenyi Lin, Qilong Wang, Pengfei Zhu, and Qinghua Hu. 2024 · 2024
Later among the works it cites.
Learn to Rectify the Bias of CLIP for Unsupervised Semantic Segmentation. In CVPR
Jingyun Wang and Guoliang Kang. 2024 · 2024
Later among the works it cites.
Visual Grounding with Multi-modal Conditional Adaptation. In ACMMM . 3877–3886
Ruilin Yao, Shengwu Xiong, Yichen Zhao, and Yi Rong. 2024 · 2024
Later among the works it cites.
Capsfusion: Rethinking image-text data at scale. In CVPR
Qiying Yu, Quan Sun, Xiaosong Zhang, Yufeng Cui, Fan Zhang, Yue Cao, Xinlong Wang, and Jingjing Liu. 2024 · 2024
Later among the works it cites.
Decoupled Global-Local Alignment for Improving Compositional Understanding
Xiaoxing Hu, Kaicheng Yang, Jun Wang, Haoran Xu, Ziyong Feng, and Yupei Wang. 2025 · 2025
Closest in time.
Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models. In ICLR
Zhengfeng Lai, Vasileios Saveris, Chen Chen, Hong-You Chen, Haotian Zhang, Bowen Zhang, Juan Lao Tebar, Wenze Hu, Zhe Gan, Peter Grasch, et al · 2025
Closest in time.
OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text. In ICLR
Qingyun Li, Zhe Chen, Weiyun Wang, Wenhai Wang, Shenglong Ye, Zhenjiang Jin, et al · 2025
Closest in time.
Scaling Pre-training to One Hundred Billion Data for Vision Language Models
Xiao Wang, Ibrahim Alabdulmohsin, Daniel Salz, Zhe Li, Keran Rong, and Xiaohua Zhai. 2025 · 2025
Closest in time.
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination. In AAAI
Kaicheng Yang, Tiancheng Gu, Xiang An, Haiqiang Jiang, Xiangzi Dai, Ziyong Feng, Weidong Cai, and Jiankang Deng. 2025 · 2025
Closest in time.