Fetching the paper…
Reading the bibliography…
Contrastive Language-Image Pre-training (CLIP) models have shown significant potential, particularly in zero-shot classification across diverse distribution shifts.
B. Schölkopf, J. Platt, and T. Hofmann, “Analysis of representations for domain adaptation,” in Advances in Neural Information Processing Systems , 2007, pp. 137–144
2007
Earlier work this paper cites.
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A Large-Scale Hierarchical Image Database,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2009
2009
Earlier work this paper cites.
J. Xiao, J. Hays, K. A. Ehinger, A. Oliva, and A. Torralba, “Sun database: Large-scale scene recognition from abbey to zoo,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2010
2010
Earlier work this paper cites.
P. J. Huber, “Robust statistics,” in International Encyclopedia of Statistical Science . Springer, 2011, pp. 1248–1251
2011
Earlier work this paper cites.
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi, “Describing textures in the wild,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2014
2014
Earlier work this paper cites.
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier, “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,” Transactions of the Association for Computational Linguistics , vol. 2, pp. 67–78, 2014. [Online]. Available: https://aclanthology.org/Q14-1006
2014
Earlier work this paper cites.
K. Nguyen and B. O’Connor, “Posterior calibration and exploratory analysis for natural language processing models,” in Conference on Empirical Methods in Natural Language Processing , 2015
2015
Earlier work this paper cites.
M. P. Naeini, G. Cooper, and M. Hauskrecht, “Obtaining well calibrated probabilities using bayesian binning,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016
2016
Earlier work this paper cites.
C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On calibration of modern neural networks,” in International Conference on Machine Learning , 2017
2017
Earlier work this paper cites.
B. Zhou, A. Lapedriza, A. Khosla, A. Oliva, and A. Torralba, “Places: A 10 million image database for scene recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 40, no. 6, pp. 1452–1464, 2017
2017
Earlier work this paper cites.
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner, “Scannet: Richly-annotated 3d reconstructions of indoor scenes,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 5828–5839
2017
Earlier work this paper cites.
D. Hendrycks and K. Gimpel, “A baseline for detecting misclassified and out-of-distribution examples in neural networks,” in International Conference on Learning Representations , 2017. [Online]. Available: https://openreview.net/forum?id=Hkg4TI9xl
2017
Earlier work this paper cites.
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2018, pp. 2556–2565
2018
Earlier work this paper cites.
Z. Wu, Y. Xiong, S. X. Yu, and D. Lin, “Unsupervised feature learning via non-parametric instance discrimination,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018
2018
Earlier work this paper cites.
G. Van Horn, O. Mac Aodha, Y. Song, Y. Cui, C. Sun, A. Shepard, H. Adam, P. Perona, and S. Belongie, “The inaturalist species classification and detection dataset,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018
2018
Earlier work this paper cites.
B. Recht, R. Roelofs, L. Schmidt, and V. Shankar, “Do imagenet classifiers generalize to imagenet?” in International Conference on Machine Learning , 2019
2019
Earlier work this paper cites.
H. Wang, S. Ge, Z. Lipton, and E. P. Xing, “Learning robust global representations by penalizing local predictive power,” in Advances in Neural Information Processing Systems , 2019
2019
Earlier work this paper cites.
A. Barbu, D. Mayo, J. Alverio, W. Luo, C. Wang, D. Gutfreund, J. Tenenbaum, and B. Katz, “Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models,” in Advances in Neural Information Processing Systems , 2019
2019
Earlier work this paper cites.
R. Geirhos, P. Rubisch, C. Michaelis, M. Bethge, F. A. Wichmann, and W. Brendel, “Imagenet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness.” in International Conference on Learning Representations , 2019. [Online]. Available: https://openreview.net/forum?id=Bygh9j09KX
2019
Earlier work this paper cites.
D. Hendrycks and T. Dietterich, “Benchmarking neural network robustness to common corruptions and perturbations,” in International Conference on Learning Representations , 2019. [Online]. Available: https://openreview.net/forum?id=HJz6tiCqYm
2019
Earlier work this paper cites.
R. Wightman, “Pytorch image models,” https://github.com/rwightman/pytorch-image-models , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” in International Conference on Machine Learning , 2019
2019
Earlier work this paper cites.
R. Geirhos, P. Rubisch, J. Rauber, C. R. M. Temme, C. Michaelis, W. Brendel, M. Bethge, and F. A. Wichmann, “Inducing a human-like shape bias leads to emergent human-level distortion robustness in cnns,” Journal of Vision , vol. 19, no. 10, pp. 209c–209c, 2019
2019
Earlier work this paper cites.
Y. Ovadia, E. Fertig, J. Ren, Z. Nado, D. Sculley, S. Nowozin, J. Dillon, B. Lakshminarayanan, and J. Snoek, “Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift,” in Advances in Neural Information Processing Systems , 2019
2019
Earlier work this paper cites.
R. Taori, A. Dave, V. Shankar, N. Carlini, B. Recht, and L. Schmidt, “Measuring robustness to natural distribution shifts in image classification,” in Advances in Neural Information Processing Systems , 2020
2020
Earlier work this paper cites.
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020
2020
Earlier work this paper cites.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in International Conference on Machine Learning , 2020
2020
Earlier work this paper cites.
R. Geirhos, J.-H. Jacobsen, C. Michaelis, R. Zemel, W. Brendel, M. Bethge, and F. A. Wichmann, “Shortcut learning in deep neural networks,” Nature Machine Intelligence , vol. 2, no. 11, pp. 665–673, 2020
2020
Earlier work this paper cites.
K. Hermann, T. Chen, and S. Kornblith, “The origins and prevalence of texture bias in convolutional neural networks,” in Advances in Neural Information Processing Systems , 2020, pp. 19 000–19 015
2020
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International Conference on Machine Learning , 2021
2021
Earlier work this paper cites.
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” in International Conference on Machine Learning , 2021
2021
Earlier work this paper cites.
D. Hendrycks, S. Basart, N. Mu, S. Kadavath, F. Wang, E. Dorundo, R. Desai, T. Zhu, S. Parajuli, M. Guo et al. , “The many faces of robustness: A critical analysis of out-of-distribution generalization,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021
2021
Earlier work this paper cites.
D. Hendrycks, K. Zhao, S. Basart, J. Steinhardt, and D. Song, “Natural adversarial examples,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021
2021
Cited alongside, same era.
J. Djolonga, J. Yung, M. Tschannen, R. Romijnders, L. Beyer, A. Kolesnikov, J. Puigcerver, M. Minderer, A. D’Amour, D. Moldovan et al. , “On robustness and transferability of convolutional neural networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 16 458–16 468
2021
Cited alongside, same era.
P. W. Koh, S. Sagawa, H. Marklund, S. M. Xie, M. Zhang, A. Balsubramani, W. Hu, M. Yasunaga, R. L. Phillips, I. Gao et al. , “Wilds: A benchmark of in-the-wild distribution shifts,” in International Conference on Machine Learning , 2021, pp. 5637–5664
2021
Cited alongside, same era.
E. Mintun, A. Kirillov, and S. Xie, “On interaction between augmentations and corruptions in natural corruption robustness,” in Advances in Neural Information Processing Systems , 2021
W. Tu, W. Deng, and T. Gedeon, “A closer look at the robustness of contrastive language-image pre-training (clip),” in Advances in Neural Information Processing Systems , 2023
2023
Later among the works it cites.
S. Y. Gadre, G. Ilharco, A. Fang, J. Hayase, G. Smyrnis, T. Nguyen, R. Marten, M. Wortsman, D. Ghosh, J. Zhang et al. , “Datacomp: In search of the next generation of multimodal datasets,” in Advances in Neural Information Processing Systems , 2023, pp. 27 092–27 112
2023
Later among the works it cites.
Z. Huang, X. Xia, L. Shen, B. Han, M. Gong, C. Gong, and T. Liu, “Harnessing out-of-distribution examples via augmenting content and style,” in International Conference on Learning Representations , 2023. [Online]. Available: https://openreview.net/forum?id=boNyg20-JDm
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=YicbFdNTTy
2021
Cited alongside, same era.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021
2021
Cited alongside, same era.
I. O. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Unterthiner, J. Yung, A. Steiner, D. Keysers, J. Uszkoreit et al. , “Mlp-mixer: An all-mlp architecture for vision,” in Advances in Neural Information Processing Systems , 2021
2021
Cited alongside, same era.
G. Ilharco, M. Wortsman, R. Wightman, C. Gordon, N. Carlini, R. Taori, A. Dave, V. Shankar, H. Namkoong, J. Miller, H. Hajishirzi, A. Farhadi, and L. Schmidt, “Openclip,” Jul. 2021, if you use this software, please cite it as below. [Online]. Available: https://doi.org/10.5281/zenodo.5143773
2021
Cited alongside, same era.
V. Shankar, A. Dave, R. Roelofs, D. Ramanan, B. Recht, and L. Schmidt, “Do image classifiers generalize across time?” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021
2021
Cited alongside, same era.
J. P. Miller, R. Taori, A. Raghunathan, S. Sagawa, P. W. Koh, V. Shankar, P. Liang, Y. Carmon, and L. Schmidt, “Accuracy on the line: on the strong correlation between out-of-distribution and in-distribution generalization,” in International Conference on Machine Learning , 2021
2021
Cited alongside, same era.
M. M. Naseer, K. Ranasinghe, S. H. Khan, M. Hayat, F. Shahbaz Khan, and M.-H. Yang, “Intriguing properties of vision transformers,” in Advances in Neural Information Processing Systems , 2021
2021
Cited alongside, same era.
I. Bello, W. Fedus, X. Du, E. D. Cubuk, A. Srinivas, T.-Y. Lin, J. Shlens, and B. Zoph, “Revisiting resnets: Improved training and scaling strategies,” in Advances in Neural Information Processing Systems , 2021
2021
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, “Sigmoid loss for language image pre-training,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 11 975–11 986
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Bitterwolf, M. Müller, and M. Hein, “In or out? Fixing ImageNet out-of-distribution detection evaluation,” in Proceedings of International Conference on Machine Learning , 2023, pp. 2471–2506
2023
Later among the works it cites.
V. Jampani, K.-K. Maninis, A. Engelhardt, A. Karpur, K. Truong, K. Sargent, S. Popov, A. Araujo, R. Martin Brualla, K. Patel et al. , “Navi: Category-agnostic image collections with high-quality 3d shape and pose annotations,” in Advances in Neural Information Processing Systems Dataset and Benchmark Track , 2023
2023
Later among the works it cites.
Y. Xiao, Z. Tang, P. Wei, C. Liu, and L. Lin, “Masked images are counterfactual samples for robust fine-tuning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 20 301–20 310
2023
Later among the works it cites.
S. Goyal, A. Kumar, S. Garg, Z. Kolter, and A. Raghunathan, “Finetune like you pretrain: Improved finetuning of zero-shot vision models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 19 338–19 347
2023
Later among the works it cites.
S. Pratt, I. Covert, R. Liu, and A. Farhadi, “What does a platypus look like? generating customized prompts for zero-shot image classification,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
M. Maniparambil, C. Vorster, D. Molloy, N. Murphy, K. McGuinness, and N. E. O’Connor, “Enhancing clip with gpt-4: Harnessing visual descriptions as prompts,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 262–271
2023
Later among the works it cites.
M. U. khattak, H. Rasheed, M. Maaz, S. Khan, and F. S. Khan, “Maple: Multi-modal prompt learning,” in The IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023
2023
Later among the works it cites.
M. U. Khattak, S. T. Wasim, M. Naseer, S. Khan, M.-H. Yang, and F. S. Khan, “Self-regulating prompts: Foundational model adaptation without forgetting,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2023, pp. 15 190–15 200
2023
Later among the works it cites.
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in International Conference on Machine Learning . PMLR, 2023, pp. 19 730–19 742
2023
Later among the works it cites.
M. El Banani, A. Raj, K.-K. Maninis, A. Kar, Y. Li, M. Rubinstein, D. Sun, L. Guibas, J. Johnson, and V. Jampani, “Probing the 3d awareness of visual foundation models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 21 795–21 806
2024
Closest in time.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in Advances in neural information processing systems , 2024
2024
Closest in time.
C. Zhang, F. Pan, J. Kim, I. S. Kweon, and C. Mao, “Imagenet-d: Benchmarking neural network robustness on diffusion synthetic object,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 21 752–21 762
2024
Closest in time.
Z. Huang, C. Liu, Y. Dong, H. Su, S. Zheng, and T. Liu, “Machine vision therapy: Multimodal large language models can enhance visual robustness via denoising in-context learning,” in International Conference on Machine Learning , 2024, pp. 19 973–20 003
2024
Closest in time.
2024
Closest in time.
A. Fang, A. M. Jose, A. Jain, L. Schmidt, A. T. Toshev, and V. Shankar, “Data filtering networks,” in The International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=KAk6ngZ09F
2024
Closest in time.
H. Xu, S. Xie, X. E. Tan, P.-Y. Huang, R. Howes, V. Sharma, S.-W. Li, G. Ghosh, L. Zettlemoyer, and C. Feichtenhofer, “Demystifying clip data,” in International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=5BCFlnfE1g
2024
Closest in time.
S. Karamcheti, S. Nair, A. Balakrishna, P. Liang, T. Kollar, and D. Sadigh, “Prismatic vlms: Investigating the design space of visually-conditioned language models,” in International Conference on Machine Learning , 2024
2024
Closest in time.
T. Li, Z. Wen, Y. Li, and T. S. Lee, “Emergence of shape bias in convolutional neural networks through activation sparsity,” in Advances in Neural Information Processing Systems , 2024
2024
Closest in time.
J. Zhang, C. Herrmann, J. Hur, E. Chen, V. Jampani, D. Sun, and M.-H. Yang, “Telling left from right: Identifying geometry-aware semantic correspondence,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 3076–3085
2024
Closest in time.
J. Zhang, C. Herrmann, J. Hur, L. Polania Cabrera, V. Jampani, D. Sun, and M.-H. Yang, “A tale of two features: Stable diffusion complements dino for zero-shot semantic correspondence,” in Advances in Neural Information Processing Systems , 2024
2024
Closest in time.
M. Oquab, T. Darcet, T. Moutakanni, H. V. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. HAZIZA, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y. Huang, S.-W. Li, I. Misra, M. Rabbat, V. Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski, “DINOv2: Learning robust visual features without supervision,” Transactions on Machine Learning Research , 2024. [Online]. Available: https://openreview.net/forum?id=a68SUt6zFt
2024
Closest in time.
O. F. Kar, A. Tonioni, P. Poklukar, A. Kulshrestha, A. Zamir, and F. Tombari, “Brave: Broadening the visual encoding of vision-language models,” in European Conference on Computer Vision , 2024, pp. 113–132
2024
Closest in time.
M. J. Mirza, L. Karlinsky, W. Lin, S. Doveh, J. Micorek, M. Kozinski, H. Kuehne, and H. Possegger, “Meta-prompting for automating zero-shot visual recognition with llms,” in European Conference on Computer Vision , 2024, pp. 370–387
2024
Closest in time.
S. Parashar, Z. Lin, T. Liu, X. Dong, Y. Li, D. Ramanan, J. Caverlee, and S. Kong, “The neglected tails in vision-language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 12 988–12 997
2024
Closest in time.
J. Chen, Q. Yu, X. Shen, A. Yuille, and L.-C. Chen, “Vitamin: Designing scalable vision models in the vision-language era,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024
2024
Closest in time.
2024
Closest in time.
C. Schuhmann and R. Beaumont, “Laion-aesthetics,” https://laion.ai/blog/laion-aesthetics/ , 2022, accessed: 2025-04-25
2025
Closest in time.