Fetching the paper…
Reading the bibliography…
Large pre-trained vision-language models, such as CLIP, have shown remarkable generalization capabilities across various tasks when appropriate text prompts are provided.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Bag-of-visual-words and spatial extensions for land-use classification
Yang, Y.; and Newsam, S. 2010 · 2010
Earlier work this paper cites.
Satellite Image Classification via Two-Layer Sparse Coding With Biased Image Representation
Dai, D.; and Yang, W. 2011 · 2011
Earlier work this paper cites.
Write a classifier: Zero-shot learning using purely textual descriptions
Elhoseiny, M.; Saleh, B.; and Elgammal, A. 2013 · 2013
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Frome, A.; Corrado, G. S.; Shlens, J.; Bengio, S.; Dean, J.; Ranzato, M.; and Mikolov, T. 2013 · 2013
Earlier work this paper cites.
Zero-shot learning through cross-modal transfer
Socher, R.; Ganjoo, M.; Manning, C. D.; and Ng, A. 2013 · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K.; and Zisserman, A. 2014 · 2014
Earlier work this paper cites.
Fast r-cnn
Girshick, R. 2015 · 2015
Earlier work this paper cites.
Predicting deep zero-shot convolutional neural networks using textual descriptions
Lei Ba, J.; Swersky, K.; Fidler, S.; et al. 2015 · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Long, J.; Shelhamer, E.; and Darrell, T. 2015 · 2015
Earlier work this paper cites.
Evaluation methods for unsupervised word embeddings
Schnabel, T.; Labutov, I.; Mimno, D.; and Joachims, T. 2015 · 2015
Earlier work this paper cites.
Deep learning based feature selection for remote sensing scene classification
Zou, Q.; Ni, L.; Zhang, T.; and Wang, Q. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Learning visual features from large weakly supervised data
Joulin, A.; Van Der Maaten, L.; Jabri, A.; and Vasilache, N. 2016 · 2016
Earlier work this paper cites.
Multi-class texture analysis in colorectal cancer histology
Kather, J.; Weis, C.; Bianconi, F.; Melchers, S.; Schad, L.; Gaiser, T.; Marx, A.; and F, Z. 2016 · 2016
Earlier work this paper cites.
You only look once: Unified, real-time object detection
Redmon, J.; Divvala, S.; Girshick, R.; and Farhadi, A. 2016 · 2016
Cited alongside, same era.
Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs
Chen, L.-C.; Papandreou, G.; Kokkinos, I.; Murphy, K.; and Yuille, A. L. 2017 · 2017
Cited alongside, same era.
Remote sensing image scene classification: Benchmark and state of the art
Cheng, G.; Han, J.; and Lu, X. 2017 · 2017
Cited alongside, same era.
Self-supervised learning of visual features through embedding images into text topic spaces
Gomez, L.; Patel, Y.; Rusinol, M.; Karatzas, D.; and Jawahar, C. 2017 · 2017
Cited alongside, same era.
Exploring models and data for remote sensing image caption generation
Lu, X.; Wang, B.; Zheng, X.; and Li, X. 2017 · 2017
Cited alongside, same era.
FILIP: fine-grained interactive language-image pre-training
Yao, L.; Huang, R.; Hou, L.; Lu, G.; Niu, M.; Xu, H.; Liang, X.; Li, Z.; Jiang, X.; and Xu, C. 2021 · 2021
Later among the works it cites.
Florence: A new foundation model for computer vision
Yuan, L.; Chen, D.; Chen, Y.-L.; Codella, N.; Dai, X.; Gao, J.; Hu, H.; Huang, X.; Li, B.; Li, C.; et al. 2021 · 2021
Later among the works it cites.
Tip-adapter: Training-free clip-adapter for better vision-language modeling
Zhang, R.; Fang, R.; Zhang, W.; Gao, P.; Li, K.; Dai, J.; Qiao, Y.; and Li, H. 2021 · 2021
Later among the works it cites.
Bridging the gap between object and image-level representations for open-vocabulary detection
Bangalath, H.; Maaz, M.; Khattak, M. U.; Khan, S. H.; and Shahbaz Khan, F. 2022 · 2022
Later among the works it cites.
Promptdet: Towards open-vocabulary detection using uncurated images
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xia, G.-S.; Hu, J.; Hu, F.; Shi, B.; Bai, X.; Zhong, Y.; Zhang, L.; and Lu, X. 2017 · 2017
Cited alongside, same era.
Frage: Frequency-agnostic word representation
Gong, C.; He, D.; Tan, X.; Qin, T.; Wang, L.; and Liu, T.-Y. 2018 · 2018
Cited alongside, same era.
PatternNet: A benchmark dataset for performance evaluation of remote sensing image retrieval
Zhou, W.; Newsam, S.; Li, C.; and Shao, Z. 2018 · 2018
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Cited alongside, same era.
MLRSNet: A multi-label high spatial resolution remote sensing dataset for semantic scene understanding
Qi, X.; Zhu, P.; Wang, Y.; Zhang, L.; Peng, J.; Wu, M.; Chen, J.; Zhao, X.; Zang, N.; and Mathiopoulos, P. T. 2020 · 2020
Cited alongside, same era.
CLIP-Art: Contrastive pre-training for fine-grained art classification
Conde, M. V.; and Turgutlu, K. 2021 · 2021
Cited alongside, same era.
Clip-adapter: Better vision-language models with feature adapters
Gao, P.; Geng, S.; Zhang, R.; Ma, T.; Fang, R.; Zhang, Y.; Li, H.; and Qiao, Y. 2021 · 2021
Cited alongside, same era.
Feng, C.; Zhong, Y.; Jie, Z.; Chu, X.; Ren, H.; Wei, X.; Xie, W.; and Ma, L. 2022 · 2022
Later among the works it cites.
Cma-clip: Cross-modality attention clip for text-image classification
Fu, J.; Xu, S.; Liu, H.; Liu, Y.; Xie, N.; Wang, C.-C.; Liu, J.; Sun, Y.; and Wang, B. 2022 · 2022
Later among the works it cites.
Domain adaptation via prompt learning
Ge, C.; Huang, R.; Xie, M.; Lai, Z.; Song, S.; Li, S.; and Huang, G. 2022 · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
He, K.; Chen, X.; Xie, S.; Li, Y.; Dollár, P.; and Girshick, R. 2022 · 2022
Later among the works it cites.
Language-driven semantic segmentation
Li, B.; Weinberger, K. Q.; Belongie, S.; Koltun, V.; and Ranftl, R. 2022 · 2022
Later among the works it cites.
Image segmentation using text and image prompts
Lüddecke, T.; and Ecker, A. 2022 · 2022
Later among the works it cites.
Denseclip: Language-guided dense prediction with context-aware prompting
Rao, Y.; Zhao, W.; Chen, G.; Tang, Y.; Zhu, Z.; Huang, G.; Zhou, J.; and Lu, J. 2022 · 2022
Later among the works it cites.
Ringmo: A remote sensing foundation model with masked image modeling
Sun, X.; Wang, P.; Lu, W.; Zhu, Z.; Lu, X.; He, Q.; Li, J.; Rong, X.; Yang, Z.; Chang, H.; et al. 2022 · 2022
Later among the works it cites.
Lit: Zero-shot transfer with locked-image text tuning
Zhai, X.; Wang, X.; Mustafa, B.; Steiner, A.; Keysers, D.; Kolesnikov, A.; and Beyer, L. 2022 · 2022
Later among the works it cites.
Crystal Clean: Brain Tumors MRI Dataset
Hashemi, S. M. H. 2023 · 2023
Closest in time.
Maple: Multi-modal prompt learning
Khattak, M. U.; Rasheed, H.; Maaz, M.; Khan, S.; and Khan, F. S. 2023 · 2023
Closest in time.
Segment anything in medical images
Ma, J.; and Wang, B. 2023 · 2023
Closest in time.