Fetching the paper…
Reading the bibliography…
Large-scale pre-trained Vision-Language Models (VLMs), such as CLIP, establish the correlation between texts and images, achieving remarkable success on various downstream tasks with fine-tuning.
Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pp. 178–178
Fei-Fei, L., Fergus, R., Perona, P., 2004 · 2004
Earlier work this paper cites.
A bayesian hierarchical model for learning natural scene categories, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 524–531
Fei-Fei, L., Perona, P., 2005 · 2005
Earlier work this paper cites.
Automated flower classification over a large number of classes, in: Indian Conference on Computer Vision, Graphics & Image Processing
Nilsback, M.E., Zisserman, A., 2008 · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 248–255
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L., 2009 · 2009
Earlier work this paper cites.
Learning to detect unseen object classes by between-class attribute transfer, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 951–958
Lampert, C.H., Nickisch, H., Harmeling, S., 2009 · 2009
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3485–3492
Xiao, J., Hays, J., Ehinger, K.A., Oliva, A., Torralba, A., 2010 · 2010
Earlier work this paper cites.
Recognizing human actions by attributes, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3337–3344
Liu, J., Kuipers, B., Savarese, S., 2011 · 2011
Earlier work this paper cites.
Cats and dogs, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3498–3505
Parkhi, O.M., Vedaldi, A., Zisserman, A., Jawahar, C., 2012 · 2012
Earlier work this paper cites.
Sun attribute database: Discovering, annotating, and recognizing scene attributes, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2751–2758
Patterson, G., Hays, J., 2012 · 2012
Earlier work this paper cites.
Attribute learning in large-scale datasets, in: European Conference on Computer Vision Workshops, pp. 1–14
Russakovsky, O., Fei-Fei, L., 2012 · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Soomro, K., Zamir, A.R., Shah, M., 2012 · 2012
Earlier work this paper cites.
3d object representations for fine-grained categorization, in: IEEE/CVF International Conference on Computer Vision Workshops, pp. 554–561
Krause, J., Stark, M., Deng, J., Fei-Fei, L., 2013 · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft
Maji, S., Rahtu, E., Kannala, J., Blaschko, M., Vedaldi, A., 2013 · 2013
Earlier work this paper cites.
Food-101–mining discriminative components with random forests, in: European Conference on Computer Vision, pp. 446–461
Bossard, L., Guillaumin, M., Van Gool, L., 2014 · 2014
Earlier work this paper cites.
Sparse representations based attribute learning for flower classification
Cheng, K., Tan, X., 2014 · 2014
Earlier work this paper cites.
Describing textures in the wild, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3606–3613
Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., Vedaldi, A., 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D.P., Ba, J., 2014 · 2014
Earlier work this paper cites.
Attribute relation learning for zero-shot classification
Liu, M., Zhang, D., Chen, S., 2014 · 2014
Earlier work this paper cites.
How to transfer? zero-shot object recognition via hierarchical transfer of semantic attributes, in: IEEE Winter Conference on Applications of Computer Vision, pp. 837–843
Al-Halah, Z., Stiefelhagen, R., 2015 · 2015
Earlier work this paper cites.
Multiview triplet embedding: Learning attributes in multiple maps, in: International Conference on Machine Learning, pp. 1472–1480
Amid, E., Ukkonen, A., 2015 · 2015
Earlier work this paper cites.
Predicting deep zero-shot convolutional neural networks using textual descriptions, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4247–4255
Lei Ba, J., Swersky, K., Fidler, S., et al., 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 770–778
He, K., Zhang, X., Ren, S., Sun, J., 2016 · 2016
Earlier work this paper cites.
Unsupervised learning of discriminative attributes and visual representations, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5175–5184
Huang, C., Loy, C.C., Tang, X., 2016 · 2016
Earlier work this paper cites.
Coco attributes: Attributes for people, animals, and objects, in: European Conference on Computer Vision, pp. 85–100
Patterson, G., Hays, J., 2016 · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks, in: International Conference on Machine Learning, PMLR. pp. 1126–1135
Finn, C., Abbeel, P., Levine, S., 2017 · 2017
Earlier work this paper cites.
Self-supervised learning of visual features through embedding images into text topic spaces, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4230–4239
Gomez, L., Patel, Y., Rusinol, M., Karatzas, D., Jawahar, C., 2017 · 2017
Earlier work this paper cites.
Low-shot learning with imprinted weights, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5822–5830
Qi, H., Brown, M., Lowe, D.G., 2018 · 2018
Earlier work this paper cites.
Attribute-aware semantic segmentation of road scenes for understanding pedestrian orientations, in: International Conference on Intelligent Transportation Systems, pp. 2698–2703
Sulistiyo, M.D., Kawanishi, Y., Deguchi, D., Hirayama, T., Ide, I., Zheng, J., Murase, H., 2018 · 2018
Earlier work this paper cites.
Deep visual domain adaptation: A survey
Wang, M., Deng, W., 2018 · 2018
Cited alongside, same era.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
Helber, P., Bischke, B., Dengel, A., Borth, D., 2019 · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for nlp, in: International Conference on Machine Learning, PMLR. pp. 2790–2799
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., Gelly, S., 2019 · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks, in: Advances in Neural Information Processing Systems
Lu, J., Batra, D., Parikh, D., Lee, S., 2019 · 2019
Cited alongside, same era.
Towards latent attribute discovery from triplet similarities, in: IEEE/CVF International Conference on Computer Vision, pp. 402–410
Nigam, I., Tokmakov, P., Ramanan, D., 2019 · 2019
Cited alongside, same era.
Pali: A jointly-scaled multilingual language-image model
Chen, X., Wang, X., Changpinyo, S., Piergiovanni, A., Padlewski, P., Salz, D., Goodman, S., Grycner, A., Mustafa, B., Beyer, L., et al., 2022 · 2022
Later among the works it cites.
Contrastive vision-language pre-training with limited resources, in: European Conference on Computer Vision, Springer. pp. 236–253
Cui, Q., Zhou, B., Guo, Y., Yin, W., Wu, H., Yoshie, O., Chen, Y., 2022 · 2022
Later among the works it cites.
Rlprompt: Optimizing discrete text prompts with reinforcement learning
Deng, M., Wang, J., Hsieh, C.P., Wang, Y., Guo, H., Shu, T., Song, M., Xing, E.P., Hu, Z., 2022 · 2022
Later among the works it cites.
Multi-modal alignment using representation codebook, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15651–15660
Duan, J., Chen, L., Tran, S., Yang, J., Xu, Y., Zeng, B., Chilimbi, T., 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Do imagenet classifiers generalize to imagenet?, in: International Conference on Machine Learning
Recht, B., Roelofs, R., Schmidt, L., Shankar, V., 2019 · 2019
Cited alongside, same era.
Learning robust global representations by penalizing local predictive power, in: Advances in Neural Information Processing Systems
Wang, H., Ge, S., Lipton, Z., Xing, E.P., 2019 · 2019
Cited alongside, same era.
A large-scale attribute dataset for zero-shot learning, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops
Zhao, B., Fu, Y., Liang, R., Wu, J., Wang, Y., Wang, Y., 2019 · 2019
Cited alongside, same era.
Improved few-shot visual classification, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14493–14502
Bateni, P., Goyal, R., Masrani, V., Wood, F., Sigal, L., 2020 · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale, in: International Conference on Learning Representations
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al., 2020 · 2020
Cited alongside, same era.
How can we know what language models know?
Jiang, Z., Xu, F.F., Araki, J., Neubig, G., 2020 · 2020
Cited alongside, same era.
Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation, in: International Conference on Machine Learning, PMLR. pp. 6028–6039
Liang, J., Hu, D., Feng, J., 2020 · 2020
Cited alongside, same era.
Huang, T., Chu, J., Wei, F., 2022 · 2022
Later among the works it cites.
Visual prompt tuning, in: European Conference on Computer Vision
Jia, M., Tang, L., Chen, B.C., Cardie, C., Belongie, S., Hariharan, B., Lim, S.N., 2022 · 2022
Later among the works it cites.
Diffusionclip: Text-guided diffusion models for robust image manipulation, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2426–2435
Kim, G., Kwon, T., Ye, J.C., 2022 · 2022
Later among the works it cites.
Grounded language-image pre-training, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10965–10975
Li, L.H., Zhang, P., Zhang, H., Yang, J., Li, C., Zhong, Y., Wang, L., Yuan, L., Zhang, L., Hwang, J.N., et al., 2022 · 2022
Later among the works it cites.
Prompt distribution learning, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5206–5215
Lu, Y., Liu, J., Zhang, Y., Liu, Y., Tian, X., 2022 · 2022
Later among the works it cites.
Test-time prompt tuning for zero-shot generalization in vision-language models, in: Advances in Neural Information Processing Systems
Manli, S., Weili, N., De-An, H., Zhiding, Y., Tom, G., Anima, A., Chaowei, X., 2022 · 2022
Later among the works it cites.
Svl-adapter: Self-supervised adapter for vision-language pretrained models, in: British Machine Vision Conference
Pantazis, O., Brostow, G., Jones, K., Mac Aodha, O., 2022 · 2022
Later among the works it cites.
Learning affinity from attention: End-to-end weakly-supervised semantic segmentation with transformers, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16846–16855
Ru, L., Zhan, Y., Yu, B., Du, B., 2022 · 2022
Later among the works it cites.
Unified vision and language prompt learning
Zang, Y., Li, W., Zhou, K., Huang, C., Loy, C.C., 2022 · 2022
Later among the works it cites.
Tip-adapter: Training-free adaption of clip for few-shot classification, in: European Conference on Computer Vision
Zhang, R., Zhang, W., Fang, R., Gao, P., Li, K., Dai, J., Qiao, Y., Li, H., 2022 · 2022
Later among the works it cites.
Prompt-aligned gradient for prompt tuning
Zhu, B., Niu, Y., Han, Y., Wu, Y., Zhang, H., 2022 · 2022
Later among the works it cites.
Prompt learning with optimal transport for vision-language models, in: International Conference on Learning Representations
Chen, G., Yao, W., Song, X., Li, X., Rao, Y., Zhang, K., 2023 · 2023
Closest in time.
HiCLIP: Contrastive language-image pretraining with hierarchy-aware attention, in: International Conference on Learning Representations
Geng, S., Yuan, J., Tian, Y., Chen, Y., Zhang, Y., 2023 · 2023
Closest in time.
Self-correctable and adaptable inference for generalizable human pose estimation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5537–5546
Kan, Z., Chen, S., Zhang, C., Tang, Y., He, Z., 2023 · 2023
Closest in time.
A comprehensive survey on test-time adaptation under distribution shifts
Liang, J., He, R., Tan, T., 2023 · 2023
Closest in time.
Multimodality helps unimodality: Cross-modal few-shot learning with multimodal models
Lin, Z., Yu, S., Kuang, Z., Pathak, D., Ramana, D., 2023 · 2023
Closest in time.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., Neubig, G., 2023 · 2023
Closest in time.
Meta learning to bridge vision and language models for multimodal few-shot learning, in: International Conference on Learning Representations
Najdenkoska, I., Zhen, X., Worring, M., 2023 · 2023
Closest in time.
Neuro-modulated hebbian learning for fully test-time adaptation
Tang, Y., Zhang, C., Xu, H., Chen, S., Cheng, J., Leng, L., Guo, Q., He, Z., 2023 · 2023
Closest in time.
Learning to decompose visual features with latent textual prompts, in: International Conference on Learning Representations
Wang, F., Li, M., Lin, X., Lv, H., Schwing, A., Ji, H., 2023 · 2023
Closest in time.
When and why vision-language models behave like bags-of-words, and what to do about it?, in: International Conference on Learning Representations
Yuksekgonul, M., Bianchi, F., Kalluri, P., Jurafsky, D., Zou, J., 2023 · 2023
Closest in time.
Domain generalization: A survey
Zhou, K., Liu, Z., Qiao, Y., Xiang, T., Loy, C.C., 2023 · 2023
Closest in time.
Attribute based object identification, in: IEEE International Conference on Robotics and Automation, pp. 2096–2103
Sun, Y., Bo, L., Fox, D., 2013 · 2096
Closest in time.