Fetching the paper…
Reading the bibliography…
Contrastive Language-Image Pre-training (CLIP) has achieved excellent performance over a wide range of tasks.
Efficiency of a good but not linear set union algorithm
Tarjan, R. E. 1975 · 1975
Earlier work this paper cites.
Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories
Fei-Fei, L.; Fergus, R.; and Perona, P. 2004 · 2004
Earlier work this paper cites.
Learning a similarity metric discriminatively, with application to face verification
Chopra, S.; Hadsell, R.; and LeCun, Y. 2005 · 2005
Earlier work this paper cites.
Automated flower classification over a large number of classes
Nilsback, M.-E.; and Zisserman, A. 2008 · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A.; Hinton, G.; et al. 2009 · 2009
Earlier work this paper cites.
Meal v2: Boosting vanilla resnet-50 to 80%+ top-1 accuracy on imagenet without tricks
Shen, Z.; and Savvides, M. 2020 · 2009
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo
Xiao, J.; Hays, J.; Ehinger, K. A.; Oliva, A.; and Torralba, A. 2010 · 2010
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning
Coates, A.; Ng, A.; and Lee, H. 2011 · 2011
Earlier work this paper cites.
Cats and dogs
Parkhi, O. M.; Vedaldi, A.; Zisserman, A.; and Jawahar, C. 2012 · 2012
Earlier work this paper cites.
Multimodal learning with deep boltzmann machines
Srivastava, N.; and Salakhutdinov, R. R. 2012 · 2012
Earlier work this paper cites.
3d object representations for fine-grained categorization
Krause, J.; Stark, M.; Deng, J.; and Fei-Fei, L. 2013 · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft
Maji, S.; Rahtu, E.; Kannala, J.; Blaschko, M.; and Vedaldi, A. 2013 · 2013
Earlier work this paper cites.
Birdsnap: Large-scale fine-grained visual categorization of birds
Berg, T.; Liu, J.; Woo Lee, S.; Alexander, M. L.; Jacobs, D. W.; and Belhumeur, P. N. 2014 · 2014
Earlier work this paper cites.
Food-101–mining discriminative components with random forests
Bossard, L.; Guillaumin, M.; and Van Gool, L. 2014 · 2014
Earlier work this paper cites.
Describing textures in the wild
Cimpoi, M.; Maji, S.; Kokkinos, I.; Mohamed, S.; and Vedaldi, A. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Euclidean distance matrices: essential theory, algorithms, and applications
Dokmanic, I.; Parhizkar, R.; Ranieri, J.; and Vetterli, M. 2015 · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G.; Vinyals, O.; and Dean, J. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Improved deep metric learning with multi-class n-pair loss objective
Sohn, K. 2016 · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Deep clustering for unsupervised learning of visual features
Caron, M.; Bojanowski, P.; Joulin, A.; and Douze, M. 2018 · 2018
Earlier work this paper cites.
Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models
Barbu, A.; Mayo, D.; Alverio, J.; Luo, W.; Wang, C.; Gutfreund, D.; Tenenbaum, J.; and Katz, B. 2019 · 2019
Cited alongside, same era.
Deep multimodal representation learning: A survey
Guo, W.; Wang, J.; and Wang, S. 2019 · 2019
Cited alongside, same era.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
Helber, P.; Bischke, B.; Dengel, A.; and Borth, D. 2019 · 2019
Cited alongside, same era.
Billion-scale similarity search with gpus
Johnson, J.; Douze, M.; and Jégou, H. 2019 · 2019
Cited alongside, same era.
Do imagenet classifiers generalize to imagenet?
Recht, B.; Roelofs, R.; Schmidt, L.; and Shankar, V. 2019 · 2019
Cited alongside, same era.
Learning Robust Global Representations by Penalizing Local Predictive Power
Transferring pre-trained multimodal representations with cross-modal similarity matching
Kim, B.; Choi, S.; Hwang, D.; Lee, M.; and Lee, H. 2022 · 2022
Later among the works it cites.
Slip: Self-supervision meets language-image pre-training
Mu, N.; Kirillov, A.; Wagner, D.; and Xie, S. 2022 · 2022
Later among the works it cites.
Unsupervised visual representation learning by online constrained k-means
Qian, Q.; Xu, Y.; Hu, J.; Li, H.; and Jin, R. 2022 · 2022
Later among the works it cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C.; Beaumont, R.; Vencu, R.; Gordon, C.; Wightman, R.; Cherti, M.; Coombes, T.; Katta, A.; Mullis, C.; Wortsman, M.; et al. 2022 · 2022
Later among the works it cites.
Clip-td: Clip targeted distillation for vision-language tasks
Wang, Z.; Codella, N.; Chen, Y.-C.; Zhou, L.; Yang, J.; Dai, X.; Xiao, B.; You, H.; Chang, S.-F.; and Yuan, L. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang, H.; Ge, S.; Lipton, Z.; and Xing, E. P. 2019 · 2019
Cited alongside, same era.
Self-labelling via simultaneous clustering and representation learning
Asano, Y. M.; Rupprecht, C.; and Vedaldi, A. 2020 · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
Chen, T.; Kornblith, S.; Norouzi, M.; and Hinton, G. 2020 · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. 2020 · 2020
Cited alongside, same era.
Contrastive multiview coding
Tian, Y.; Krishnan, D.; and Isola, P. 2020 · 2020
Cited alongside, same era.
Cm-bert: Cross-modal bert for text-audio sentiment analysis
Yang, K.; Xu, H.; and Gao, K. 2020 · 2020
Cited alongside, same era.
Ch-sims: A chinese multimodal sentiment analysis dataset with fine-grained annotation of modality
Yu, W.; Xu, H.; Meng, F.; Zhu, Y.; Ma, Y.; Wu, J.; Zou, J.; and Yang, K. 2020 · 2020
Cited alongside, same era.
Contrastive learning rivals masked image modeling in fine-tuning via feature distillation
Wei, Y.; Hu, H.; Xie, Z.; Zhang, Z.; Cao, Y.; Bao, J.; Chen, D.; and Guo, B. 2022 · 2022
Later among the works it cites.
SemDeDup: Data-efficient learning at web-scale through semantic deduplication
Abbas, A.; Tirumala, K.; Simig, D.; Ganguli, S.; and Morcos, A. S. 2023 · 2023
Later among the works it cites.
Unicom: Universal and Compact Representation Learning for Image Retrieval
An, X.; Deng, J.; Yang, K.; Li, J.; Feng, Z.; Guo, J.; Yang, J.; and Liu, T. 2023 · 2023
Later among the works it cites.
Less is More: Removing Text-regions Improves CLIP Training Efficiency and Robustness
Cao, L.; Zhang, B.; Chen, C.; Yang, Y.; Du, X.; Zhang, W.; Lu, Z.; and Zheng, Y. 2023 · 2023
Later among the works it cites.
Exploring open-vocabulary semantic segmentation from clip vision encoder distillation only
Chen, J.; Zhu, D.; Qian, G.; Ghanem, B.; Yan, Z.; Zhu, C.; Xiao, F.; Culatana, S. C.; and Elhoseiny, M. 2023 · 2023
Later among the works it cites.
HiCLIP: Contrastive Language-Image Pretraining with Hierarchy-aware Attention
Geng, S.; Yuan, J.; Tian, Y.; Chen, Y.; and Zhang, Y. 2023 · 2023
Later among the works it cites.
Scaling language-image pre-training via masking
Li, Y.; Fan, H.; Hu, R.; Feichtenhofer, C.; and He, K. 2023 · 2023
Later among the works it cites.
Filtering, distillation, and hard negatives for vision-language pre-training
Radenovic, F.; Dubey, A.; Kadian, A.; Mihaylov, T.; Vandenhende, S.; Patel, Y.; Wen, Y.; Ramanathan, V.; and Mahajan, D. 2023 · 2023
Later among the works it cites.
Hybrid Distillation: Connecting Masked Autoencoders with Contrastive Learners
Shi, B.; Zhang, X.; Wang, Y.; Li, J.; Dai, W.; Zou, J.; Xiong, H.; and Tian, Q. 2023 · 2023
Later among the works it cites.
DIME-FM: DIstilling Multimodal and Efficient Foundation Models
Sun, X.; Zhang, P.; Zhang, P.; Shah, H.; Saenko, K.; and Xia, X. 2023 · 2023
Later among the works it cites.
Too Large; Data Reduction for Vision-Language Pre-Training
Wang, A. J.; Lin, K. Q.; Zhang, D. J.; Lei, S. W.; and Shou, M. Z. 2023 · 2023
Later among the works it cites.
Tinyclip: Clip distillation via affinity mimicking and weight inheritance
Wu, K.; Peng, H.; Zhou, Z.; Xiao, B.; Liu, M.; Yuan, L.; Xuan, H.; Valenzuela, M.; Chen, X. S.; Wang, X.; et al. 2023 · 2023
Later among the works it cites.
Alip: Adaptive language-image pre-training with synthetic caption
Yang, K.; Deng, J.; An, X.; Li, J.; Feng, Z.; Guo, J.; Yang, J.; and Liu, T. 2023 · 2023
Later among the works it cites.
Effective pruning of web-scale datasets based on complexity of concept clusters
Abbas, A.; Rusak, E.; Tirumala, K.; Brendel, W.; Chaudhuri, K.; and Morcos, A. S. 2024 · 2024
Closest in time.
Multi-label Cluster Discrimination for Visual Representation Learning
An, X.; Yang, K.; Dai, X.; Feng, Z.; and Deng, J. 2024 · 2024
Closest in time.
Datacomp: In search of the next generation of multimodal datasets
Gadre, S. Y.; Ilharco, G.; Fang, A.; Hayase, J.; Smyrnis, G.; Nguyen, T.; Marten, R.; Wortsman, M.; Ghosh, D.; Zhang, J.; et al. 2024 · 2024
Closest in time.
RWKV-CLIP: A Robust Vision-Language Representation Learner
Gu, T.; Yang, K.; An, X.; Feng, Z.; Liu, D.; Cai, W.; and Deng, J. 2024 · 2024
Closest in time.