Fetching the paper…
Reading the bibliography…
Visual prompt tuning (VPT) is a promising solution incorporating learnable prompt tokens to customize pre-trained models for downstream tasks.
Learning methods for generic object recognition with invariance to pose and lighting
LeCun, Y., Huang, F. J., and Bottou, L · 2004
Earlier work this paper cites.
One-shot learning of object categories
Fei-Fei, L., Fergus, R., and Perona, P · 2006
Earlier work this paper cites.
Automated flower classification over a large number of classes
Nilsback, M.-E. and Zisserman, A · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Normalized mutual information feature selection
Estévez, P. A., Tesmer, M., Perez, C. A., and Zurada, J. M · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo
Xiao, J., Hays, J., Ehinger, K. A., Oliva, A., and Torralba, A · 2010
Earlier work this paper cites.
Novel dataset for fine-grained image categorization: Stanford dogs
Khosla, A., Jayadevaprakash, N., Yao, B., and Li, F.-F · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y · 2011
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
Wah, C., Branson, S., Welinder, P., Perona, P., and Belongie, S · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Cats and dogs
Parkhi, O. M., Vedaldi, A., Zisserman, A., and Jawahar, C · 2012
Earlier work this paper cites.
Vision meets robotics: The kitti dataset
Geiger, A., Lenz, P., Stiller, C., and Urtasun, R · 2013
Earlier work this paper cites.
Describing textures in the wild
Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., and Vedaldi, A · 2014
Earlier work this paper cites.
Kaggle diabetic retinopathy detection competition report
Graham, B · 2015
Earlier work this paper cites.
Building a bird recognition app and large scale dataset with citizen scientists: The fine print in fine-grained dataset collection
Van Horn, G., Branson, S., Farrell, R., Haber, S., Barry, J., Ipeirotis, P., Perona, P., and Belongie, S · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Beattie, C., Leibo, J. Z., Teplyashin, D., Ward, T., Wainwright, M., Küttler, H., Lefrancq, A., Green, S., Valdés, V., Sadik, A., et al · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Sgdr: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2016
Cited alongside, same era.
Remote sensing image scene classification: Benchmark and state of the art
Cheng, G., Han, J., and Lu, X · 2017
Cited alongside, same era.
Fine-grained car detection for visual census estimation
Gebru, T., Krause, J., Wang, Y., Chen, D., Deng, J., and Fei-Fei, L · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Cited alongside, same era.
Big transfer (bit): General visual representation learning
Kolesnikov, A., Beyer, L., Zhai, X., Puigcerver, J., Yung, J., Gelly, S., and Houlsby, N · 2020
Later among the works it cites.
An empirical study of training self-supervised vision transformers
Chen, X., Xie, S., and He, K · 2021
Later among the works it cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Later among the works it cites.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson, J., Hariharan, B., Van Der Maaten, L., Fei-Fei, L., Lawrence Zitnick, C., and Girshick, R · 2017
Cited alongside, same era.
dsprites: Disentanglement testing sprites dataset, 2017
Matthey, L., Higgins, I., Hassabis, D., and Lerchner, A · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Rotation equivariant cnns for digital pathology
Veeling, B. S., Linmans, J., Winkens, J., Cohen, T., and Welling, M · 2018
Cited alongside, same era.
Spottune: transfer learning through adaptive fine-tuning
Guo, Y., Shi, H., Kumar, A., Grauman, K., Rosing, T., and Feris, R · 2019
Cited alongside, same era.
Rethinking imagenet pre-training
He, K., Girshick, R., and Dollár, P · 2019
Cited alongside, same era.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
Helber, P., Bischke, B., Dengel, A., and Borth, D · 2019
Cited alongside, same era.
Li, X. L. and Liang, P · 2021
Later among the works it cites.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Zaken, E. B., Ravfogel, S., and Goldberg, Y · 2021
Later among the works it cites.
Adaptformer: Adapting vision transformers for scalable visual recognition
Chen, S., Ge, C., Tong, Z., Wang, J., Song, Y., Wang, J., and Luo, P · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2022
Later among the works it cites.
Visual prompt tuning
Jia, M., Tang, L., Chen, B.-C., Cardie, C., Belongie, S., Hariharan, B., and Lim, S.-N · 2022
Later among the works it cites.
Scaling & shifting your features: A new baseline for efficient model tuning
Lian, D., Zhou, D., Feng, J., and Wang, X · 2022
Later among the works it cites.
P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks
Liu, X., Ji, K., Fu, Y., Tam, W. L., Du, Z., Yang, Z., and Tang, J · 2022
Later among the works it cites.
A large-scale study of representation learning with the visual task adaptation benchmark
Zhai, X., Puigcerver, J., Kolesnikov, A., Ruyssen, P., Riquelme, C., Lucic, M., Djolonga, J., Pinto, A. S., Neumann, M., Dosovitskiy, A., et al · 2022
Later among the works it cites.
E2vpt: An effective and efficient approach for visual prompt tuning
Cheng, H., Qifan, W., Yiming, C., Zhiwen, C., Wenguan, W., Siyuan, Q., and Dongfang, L · 2023
Later among the works it cites.
Scaling vision transformers to 22 billion parameters
Dehghani, M., Djolonga, J., Mustafa, B., and et al · 2023
Later among the works it cites.
Tuning pre-trained model via moment probing
Gao, M., Wang, Q., Lin, Z., Zhu, P., Hu, Q., and Zhou, J · 2023
Later among the works it cites.
Visual query tuning: Towards effective usage of intermediate representations for parameter and memory efficient transfer learning
Tu, C.-H., Mai, Z., and Chao, W.-L · 2023
Later among the works it cites.
Adapting shortcut with normalizing flow: An efficient tuning framework for visual recognition
Wang, Y., Shi, B., Zhang, X., Li, J., Liu, Y., Dai, W., Li, C., Xiong, H., and Tian, Q · 2023
Later among the works it cites.
Improving visual prompt tuning for self-supervised vision transformers
Yoo, S., Kim, E., Jung, D., Lee, J., and Yoon, S · 2023
Later among the works it cites.