Fetching the paper…
Reading the bibliography…
Contrastive Language-Image Pre-trained (CLIP) models have zero-shot ability of classifying an image belonging to "[CLASS]" by using similarity between the image and the prompt sentence "a [CONTEXT] of [CLASS]".
Hadsell R, Chopra S, LeCun Y (2006) Dimensionality reduction by learning an invariant mapping. In: 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’06), IEEE, pp 1735–1742
2006
Earlier work this paper cites.
Nilsback ME, Zisserman A (2008) Automated flower classification over a large number of classes. In: ICVGIP
2008
Earlier work this paper cites.
Deng J, Dong W, Socher R, et al (2009) Imagenet: A large-scale hierarchical image database. In: CVPR
2009
Earlier work this paper cites.
Torralba A, Efros A (2011) Unbiased look at dataset bias. In: Proceedings of the 2011 IEEE Conference on Computer Vision and Pattern Recognition, pp 1521–1528
2011
Earlier work this paper cites.
Moreno-Torres JG, Raeder T, Alaiz-Rodríguez R, et al (2012) A unifying view on dataset shift in classification. Pattern recognition 45(1):521–530
2012
Earlier work this paper cites.
Parkhi OM, Vedaldi A, Zisserman A, et al (2012) Cats and dogs. In: CVPR
2012
Earlier work this paper cites.
Krause J, Stark M, Deng J, et al (2013) 3d object representations for fine-grained categorization. In: ICCV-W
2013
Earlier work this paper cites.
Maji S, Rahtu E, Kannala J, et al (2013) Fine-grained visual classification of aircraft. arXiv preprint arXiv:13065151
2013
Earlier work this paper cites.
Cimpoi M, Maji S, Kokkinos I, et al (2014) Describing textures in the wild. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 3606–3613
2014
Earlier work this paper cites.
He K, Zhang X, Ren S, et al (2016) Deep residual learning for image recognition. In: CVPR
2016
Earlier work this paper cites.
Thomee B, Shamma DA, Friedland G, et al (2016) Yfcc100m: The new data in multimedia research. Communications of the ACM 59(2):64–73
2016
Earlier work this paper cites.
Xiao J, Ehinger KA, Hays J, et al (2016) Sun database: Exploring a large collection of scene categories. International Journal of Computer Vision 119(1):3–22
2016
Earlier work this paper cites.
Ge W, Yu Y (2017) Borrowing treasures from the wealthy: Deep transfer learning through selective joint fine-tuning. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 1086–1095
2017
Earlier work this paper cites.
Krizhevsky A, Sutskever I, Hinton GE (2017) Imagenet classification with deep convolutional neural networks. Communications of the ACM 60(6):84–90
2017
Earlier work this paper cites.
Li D, Yang Y, Song YZ, et al (2017) Deeper, broader and artier domain generalization. In: Proceedings of the IEEE international conference on computer vision, pp 5542–5550
2017
Earlier work this paper cites.
Venkateswara H, Eusebio J, Chakraborty S, et al (2017) Deep hashing network for unsupervised domain adaptation. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 5018–5027
2017
Earlier work this paper cites.
Beery S, Van Horn G, Perona P (2018) Recognition in terra incognita. In: Proceedings of the European conference on computer vision (ECCV), pp 456–473
2018
Earlier work this paper cites.
Loshchilov I, Hutter F (2018) Decoupled weight decay regularization. In: International Conference on Learning Representations
2018
Earlier work this paper cites.
Barbu A, Mayo D, Alverio J, et al (2019) Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models. Advances in neural information processing systems 32
2019
Earlier work this paper cites.
Guo Y, Shi H, Kumar A, et al (2019) Spottune: transfer learning through adaptive fine-tuning. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 4805–4814
2019
Cited alongside, same era.
Hendrycks D, Mu N, Cubuk ED, et al (2019) Augmix: A simple data processing method to improve robustness and uncertainty. In: International Conference on Learning Representations
2019
Cited alongside, same era.
Peng X, Bai Q, Xia X, et al (2019) Moment matching for multi-source domain adaptation. In: Proceedings of the IEEE/CVF international conference on computer vision, pp 1406–1415
2019
Cited alongside, same era.
Petroni F, Rocktäschel T, Lewis P, et al (2019) Language models as knowledge bases? In: EMNLP
2019
Cited alongside, same era.
Recht B, Roelofs R, Schmidt L, et al (2019) Do imagenet classifiers generalize to imagenet? In: ICML
Kumar A, Raghunathan A, Jones RM, et al (2021) Fine-tuning can distort pretrained features and underperform out-of-distribution. In: International Conference on Learning Representations
2021
Later among the works it cites.
Li Y, Liang F, Zhao L, et al (2021) Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm. In: International Conference on Learning Representations
2021
Later among the works it cites.
Liu P, Yuan W, Fu J, et al (2021) Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:210713586
2021
Later among the works it cites.
Miller JP, Taori R, Raghunathan A, et al (2021) Accuracy on the line: on the strong correlation between out-of-distribution and in-distribution generalization. In: International Conference on Machine Learning, PMLR, pp 7721–7735
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
Wang H, Ge S, Lipton Z, et al (2019) Learning robust global representations by penalizing local predictive power. In: NeurIPS
2019
Cited alongside, same era.
Foret P, Kleiner A, Mobahi H, et al (2020) Sharpness-aware minimization for efficiently improving generalization. In: International Conference on Learning Representations
2020
Cited alongside, same era.
Gulrajani I, Lopez-Paz D (2020) In search of lost domain generalization. In: International Conference on Learning Representations
2020
Cited alongside, same era.
Radosavovic I, Kosaraju RP, Girshick R, et al (2020) Designing network design spaces. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 10,428–10,436
2020
Cited alongside, same era.
Taori R, Dave A, Shankar V, et al (2020) Measuring robustness to natural distribution shifts in image classification. In: NeurIPS
2020
Cited alongside, same era.
Xie C, Tan M, Gong B, et al (2020) Adversarial examples improve image recognition. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 819–828
2020
Cited alongside, same era.
Andreassen A, Bahri Y, Neyshabur B, et al (2021) The evolution of out-of-distribution robustness throughout fine-tuning. arXiv preprint arXiv:210615831
2021
Cited alongside, same era.
Pham H, Dai Z, Ghiasi G, et al (2021) Combined scaling for open-vocabulary image classification. arXiv preprint arXiv: 211110050
2021
Later among the works it cites.
Radford A, Kim JW, Hallacy C, et al (2021) Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning, PMLR, pp 8748–8763
2021
Later among the works it cites.
Yao L, Huang R, Hou L, et al (2021) Filip: Fine-grained interactive language-image pre-training. In: International Conference on Learning Representations
2021
Later among the works it cites.
Yuan L, Chen D, Chen YL, et al (2021) Florence: A new foundation model for computer vision. arXiv preprint arXiv:211111432
2021
Later among the works it cites.
Cha J, Lee K, Park S, et al (2022) Domain generalization by mutual-information regularization with pre-trained models. arXiv preprint arXiv:220310789
2022
Closest in time.
Fang A, Ilharco G, Wortsman M, et al (2022) Data determines distributional robustness in contrastive language image pre-training (clip). arXiv preprint arXiv:220501397
2022
Closest in time.
Herrmann C, Sargent K, Jiang L, et al (2022) Pyramid adversarial training improves vit performance. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 13,419–13,429
2022
Closest in time.
Li C, Liu H, Li LH, et al (2022) Elevater: A benchmark and toolkit for evaluating language-augmented visual models. arXiv preprint arXiv:220408790
2022
Closest in time.
Mu N, Kirillov A, Wagner D, et al (2022) Slip: Self-supervision meets language-image pre-training. In: European Conference on Computer Vision, Springer, pp 529–544
2022
Closest in time.
Schuhmann C, Beaumont R, Vencu R, et al (2022) Laion-5b: An open large-scale dataset for training next generation image-text models. In: Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track
2022
Closest in time.
Wang Z, Bai Y, Zhou Y, et al (2022) Can cnns be more robust than transformers? arXiv preprint arXiv:220603452
2022
Closest in time.
Yu J, Wang Z, Vasudevan V, et al (2022) Coca: Contrastive captioners are image-text foundation models. arXiv preprint arXiv:220501917
2022
Closest in time.
Zhai X, Wang X, Mustafa B, et al (2022) Lit: Zero-shot transfer with locked-image text tuning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 18,123–18,133
2022
Closest in time.
Paul S, Chen PY (2022) Vision transformers are robust learners. In: Proceedings of the AAAI Conference on Artificial Intelligence, pp 2071–2081
2081
Closest in time.