Fetching the paper…
Reading the bibliography…
Vision-language models (VLMs) have demonstrated exceptional generalization capabilities for downstream tasks.
L. Fei-Fei, R. Fergus, and P. Perona, “Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories,” in
2004
Earlier work this paper cites.
M.-E. Nilsback and A. Zisserman, “Automated flower classification over a large number of classes,” in
2008
Earlier work this paper cites.
M.-E. Nilsback and A. Zisserman, “Automated flower classification over a large number of classes,” in
2008
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in
2009
Earlier work this paper cites.
J. Xiao, J. Hays, K. A. Ehinger, A. Oliva, and A. Torralba, “Sun database: Large-scale scene recognition from abbey to zoo,” in
2010
Earlier work this paper cites.
K. Soomro, A. R. Zamir, and M. Shah, “A dataset of 101 human action classes from videos in the wild,”
2012
Earlier work this paper cites.
O. M. Parkhi, A. Vedaldi, A. Zisserman, and C. Jawahar, “Cats and dogs,” in
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
J. Krause, M. Stark, J. Deng, and L. Fei-Fei, “3d object representations for fine-grained categorization,” in
2013
Earlier work this paper cites.
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi, “Describing textures in the wild,” in
2014
Earlier work this paper cites.
L. Bossard, M. Guillaumin, and L. Van Gool, “Food-101–mining discriminative components with random forests,” in
2014
Earlier work this paper cites.
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,”
2015
Earlier work this paper cites.
N. Passalis and A. Tefas, “Learning deep representations with probabilistic knowledge transfer,” in
2018
Earlier work this paper cites.
B. Recht, R. Roelofs, L. Schmidt, and V. Shankar, “Do imagenet classifiers generalize to imagenet?,” in
2019
Earlier work this paper cites.
H. Wang, S. Ge, Z. Lipton, and E. P. Xing, “Learning robust global representations by penalizing local predictive power,”
2019
Earlier work this paper cites.
P. Helber, B. Bischke, A. Dengel, and D. Borth, “Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification,”
2019
Earlier work this paper cites.
S. I. Mirzadeh, M. Farajtabar, A. Li, N. Levine, A. Matsukawa, and H. Ghasemzadeh, “Improved knowledge distillation via teacher assistant,” in
2020
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark,
2021
Cited alongside, same era.
D. Chen, J.-P. Mei, Y. Zhang, C. Wang, Z. Wang, Y. Feng, and C. Chen, “Cross-layer distillation with semantic calibration,” in
2021
Cited alongside, same era.
D. Hendrycks, K. Zhao, S. Basart, J. Steinhardt, and D. Song, “Natural adversarial examples,” in
2021
Cited alongside, same era.
D. Hendrycks, S. Basart, N. Mu, S. Kadavath, F. Wang, E. Dorundo, R. Desai, T. Zhu, S. Parajuli, M. Guo,
2021
Cited alongside, same era.
E. Cho, J. Kim, and H. J. Kim, “Distribution-aware prompt tuning for vision-language models,” in
2023
Later among the works it cites.
W. Zhou and Z. Zhou, “Unsupervised domain adaption harnessing vision-language pre-training,”
2024
Closest in time.
H. Yao, R. Zhang, and Xu, “Tcp: Textual-based class-aware prompt tuning for visual-language model,” in
2024
Closest in time.
L. Yang, R.-Y. Zhang, Y. Wang, and X. Xie, “Mma: Multi-modal adapter for vision-language models,” in
2024
Closest in time.
W. Zhang, Y. Huang, W. Zhang, T. Zhang, Q. Lao, Y. Yu, W.-S. Zheng, and R. Wang, “Continual learning of image classes with language guidance from a vision-language model,”
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in
2022
Cited alongside, same era.
J. Ding, N. Xue, G.-S. Xia, and D. Dai, “Decoupling zero-shot semantic segmentation,” in
2022
Cited alongside, same era.
K. Zhou, J. Yang, C. C. Loy, and Z. Liu, “Learning to prompt for vision-language models,”
2022
Cited alongside, same era.
K. Zhou, J. Yang, C. C. Loy, and Z. Liu, “Conditional prompt learning for vision-language models,” in
2022
Cited alongside, same era.
M. U. Khattak, H. Rasheed, M. Maaz, S. Khan, and F. S. Khan, “Maple: Multi-modal prompt learning,” in
2023
Cited alongside, same era.
H. Yao, R. Zhang, and C. Xu, “Visual-language prompt tuning with knowledge-guided context optimization,” in
2023
Cited alongside, same era.
H. Rasheed, M. U. Khattak, M. Maaz, S. Khan, and F. S. Khan, “Fine-tuned clip models are efficient video learners,” in
2023
Cited alongside, same era.
2024
Closest in time.
Y. Wang, X. Jiang, D. Cheng, D. Li, and C. Zhao, “Learning hierarchical prompt with structured linguistic knowledge for vision-language models,” in
2024
Closest in time.
Y. Zhang, C. Zhang, K. Yu, Y. Tang, and Z. He, “Concept-guided prompt learning for generalization in vision-language models,” in
2024
Closest in time.
2024
Closest in time.
Z. Li, X. Li, X. Fu, X. Zhang, W. Wang, S. Chen, and J. Yang, “Promptkd: Unsupervised prompt distillation for vision-language models,” in
2024
Closest in time.
J. Zhang, S. Wu, L. Gao, H. T. Shen, and J. Song, “Dept: Decoupled prompt tuning,” in
2024
Closest in time.
M. Farina, M. Mancini, G. Iacca, and E. Ricci, “Rethinking few-shot adaptation of vision-language models in two stages,” in
2025
Closest in time.
P. Lu, X. Li, R. Zhu, Z. Ma, J. Cao, and J.-H. Xue, “Fine-tuning via linked domains: A closed-form dual alignment mechanism for transferring vision-language models,”
2026
Closest in time.
G. Wang, P. Zhao, X. Wang, H. Guo, N. Qi, S. Yang, and Q. Guo, “Dhpt: Dual-modality heterogeneous prompt tuning for online test-time adaption in vision-language models,”
2026
Closest in time.
Y. Liang, X. Li, X. Chen, H. Chen, Y. Zhen, Z. Liu, R. Zhu, B. Li, and X. Xue, “Pyramid token pruning for high-resolution large vision-language models via region, token, and instruction-guided importance,”
2026
Closest in time.