Fetching the paper…
Reading the bibliography…
Adapter-style efficient transfer learning (ETL) has shown excellent performance in the tuning of vision-language models (VLMs) under the low-data regime, where only a few additional parameters are introduced to excavate the task-specific knowledge based on the general and powerful representation of VLMs.
Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories
L. Fei-Fei, R. Fergus, and P. Perona · 2004
Earlier work this paper cites.
Automated flower classification over a large number of classes
M.-E. Nilsback and A. Zisserman · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo
J. Xiao, J. Hays, K. A. Ehinger, A. Oliva, and A. Torralba · 2010
Earlier work this paper cites.
Cats and dogs
O. M. Parkhi, A. Vedaldi, A. Zisserman, and C. Jawahar · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Translating embeddings for modeling multi-relational data
A. Bordes, N. Usunier, A. Garcia-Duran, J. Weston, and O. Yakhnenko · 2013
Earlier work this paper cites.
3d object representations for fine-grained categorization
J. Krause, M. Stark, J. Deng, and L. Fei-Fei · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft
S. Maji, E. Rahtu, J. Kannala, M. Blaschko, and A. Vedaldi · 2013
Earlier work this paper cites.
Food-101–mining discriminative components with random forests
L. Bossard, M. Guillaumin, and L. Van Gool · 2014
Earlier work this paper cites.
Describing textures in the wild
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
T. N. Kipf and M. Welling · 2016
Earlier work this paper cites.
Stochastic graph as a model for social networks
A. Rezvanian and M. R. Meybodi · 2016
Earlier work this paper cites.
One-shot relational learning for knowledge graphs
W. Xiong, M. Yu, S. Chang, X. Guo, and W. Y. Wang · 2018
Earlier work this paper cites.
Learning semantic-specific graph representation for multi-label image recognition
T. Chen, M. Xu, X. Hui, H. Wu, and L. Lin · 2019
Earlier work this paper cites.
Knowledge-embedded routing network for scene graph generation
T. Chen, W. Yu, R. Chen, and L. Lin · 2019
Earlier work this paper cites.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
P. Helber, B. Bischke, A. Dengel, and D. Borth · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly · 2019
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
L. H. Li, M. Yatskar, D. Yin, C.-J. Hsieh, and K.-W. Chang · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
J. Lu, D. Batra, D. Parikh, and S. Lee · 2019
Earlier work this paper cites.
Do imagenet classifiers generalize to imagenet?
B. Recht, R. Roelofs, L. Schmidt, and V. Shankar · 2019
Earlier work this paper cites.
Vl-bert: Pre-training of generic visual-linguistic representations
W. Su, X. Zhu, Y. Cao, B. Li, L. Lu, F. Wei, and J. Dai · 2019
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
H. Tan and M. Bansal · 2019
Earlier work this paper cites.
Learning robust global representations by penalizing local predictive power
H. Wang, S. Ge, Z. Lipton, and E. P. Xing · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Knowledge graph transfer network for few-shot recognition
R. Chen, T. Chen, X. Hui, H. Wu, G. Li, and L. Lin · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin · 2021
Cited alongside, same era.
Prompt distribution learning
Y. Lu, J. Liu, Y. Zhang, Y. Liu, and X. Tian · 2022
Later among the works it cites.
A survey on visual transfer learning using knowledge graphs
S. Monka, L. Halilaj, and A. Rettinger · 2022
Later among the works it cites.
Svl-adapter: Self-supervised adapter for vision-language pretrained models
O. Pantazis, G. Brostow, K. Jones, and O. Mac Aodha · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Later among the works it cites.
Denseclip: Language-guided dense prediction with context-aware prompting
Y. Rao, W. Zhao, G. Chen, Y. Tang, Z. Zhu, G. Huang, J. Zhou, and J. Lu · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Du, C. Yuan, R. Barton, T. Neiman, and H. Tong · 2021
Cited alongside, same era.
Clip-adapter: Better vision-language models with feature adapters
P. Gao, S. Geng, R. Zhang, T. Ma, R. Fang, Y. Zhang, H. Li, and Y. Qiao · 2021
Cited alongside, same era.
The many faces of robustness: A critical analysis of out-of-distribution generalization
D. Hendrycks, S. Basart, N. Mu, S. Kadavath, F. Wang, E. Dorundo, R. Desai, T. Zhu, S. Parajuli, M. Guo, et al · 2021
Cited alongside, same era.
Natural adversarial examples
D. Hendrycks, K. Zhao, S. Basart, J. Steinhardt, and D. Song · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig · 2021
Cited alongside, same era.
Graph learning regularization and transfer learning for few-shot event detection
V. D. Lai, M. V. Nguyen, T. H. Nguyen, and F. Dernoncourt · 2021
Cited alongside, same era.
Graph signal processing, graph neural network and graph learning on biological data: a systematic review
R. Li, X. Yuan, M. Radfar, P. Marendy, W. Ni, T. J. O’Brien, and P. M. Casillas-Espinosa · 2021
Cited alongside, same era.
Few-shot real image restoration via distortion-relation guided transfer learning
X. Li, X. Jin, J. Fu, X. Yu, B. Tong, and Z. Chen · 2021
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Later among the works it cites.
Test-time prompt tuning for zero-shot generalization in vision-language models
M. Shu, W. Nie, D.-A. Huang, Z. Yu, T. Goldstein, A. Anandkumar, and C. Xiao · 2022
Later among the works it cites.
Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks
Y.-L. Sung, J. Cho, and M. Bansal · 2022
Later among the works it cites.
Motionclip: Exposing human motion generation to clip space
G. Tevet, B. Gordon, A. Hertz, A. H. Bermano, and D. Cohen-Or · 2022
Later among the works it cites.
Hierarchical relational learning for few-shot knowledge graph completion
H. Wu, J. Yin, B. Rajaratnam, and J. Guo · 2022
Later among the works it cites.
Dmh-fsl: Dual-modal hypergraph for few-shot learning
R. Xu, B. Liu, X. Lu, K. Zhang, and W. Liu · 2022
Later among the works it cites.
Task residual for tuning vision-language models
T. Yu, Z. Lu, X. Jin, Z. Chen, and X. Wang · 2022
Later among the works it cites.
Tip-adapter: Training-free adaption of clip for few-shot classification
R. Zhang, W. Zhang, R. Fang, P. Gao, K. Li, J. Dai, Y. Qiao, and H. Li · 2022
Later among the works it cites.
Prompting through prototype: A prototype-based prompt learning on pretrained vision-language models
Y. Zhang, H. Fei, D. Li, T. Yu, and P. Li · 2022
Later among the works it cites.
Conditional prompt learning for vision-language models
K. Zhou, J. Yang, C. C. Loy, and Z. Liu · 2022
Later among the works it cites.
Learning to prompt for vision-language models
K. Zhou, J. Yang, C. C. Loy, and Z. Liu · 2022
Later among the works it cites.
Debiasing vision-language models via biased prompts
C.-Y. Chuang, V. Jampani, Y. Li, A. Torralba, and S. Jegelka · 2023
Closest in time.
Ablating concepts in text-to-image diffusion models
N. Kumari, B. Zhang, S.-Y. Wang, E. Shechtman, R. Zhang, and J.-Y. Zhu · 2023
Closest in time.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
P. Liu, W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig · 2023
Closest in time.
Prediction calibration for generalized few-shot semantic segmentation
Z. Lu, S. He, D. Li, Y.-Z. Song, and T. Xiang · 2023
Closest in time.
Continual diffusion: Continual customization of text-to-image diffusion with c-lora
J. Seale Smith, Y.-C. Hsu, L. Zhang, T. Hua, Z. Kira, Y. Shen, and H. Jin · 2023
Closest in time.
Clip-guided prototype modulating for few-shot action recognition
X. Wang, S. Zhang, J. Cen, C. Gao, Y. Zhang, D. Zhao, and N. Sang · 2023
Closest in time.
Vita-clip: Video and text adaptive clip via multimodal prompting
S. T. Wasim, M. Naseer, S. Khan, F. S. Khan, and M. Shah · 2023
Closest in time.
Vision-language models for vision tasks: A survey
J. Zhang, J. Huang, S. Jin, and S. Lu · 2023
Closest in time.
Prompt, generate, then cache: Cascade of foundation models makes strong few-shot learners
R. Zhang, X. Hu, B. Li, S. Huang, H. Deng, H. Li, Y. Qiao, and P. Gao · 2023
Closest in time.
Not all features matter: Enhancing few-shot clip with adaptive prior refinement
X. Zhu, R. Zhang, B. He, A. Zhou, D. Wang, B. Zhao, and P. Gao · 2023
Closest in time.