Fetching the paper…
Reading the bibliography…
Large-scale contrastive vision-language pre-training has shown significant progress in visual representation learning.
Fei-Fei L, Fergus R, Perona P (2004) Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories. In: 2004 conference on computer vision and pattern recognition workshop, IEEE, pp 178–178
2004
Earlier work this paper cites.
Van der Maaten L, Hinton G (2008) Visualizing data using t-sne. Journal of machine learning research 9(11)
2008
Earlier work this paper cites.
Nilsback ME, Zisserman A (2008) Automated flower classification over a large number of classes. In: 2008 Sixth Indian Conference on Computer Vision, Graphics & Image Processing, IEEE, pp 722–729
2008
Earlier work this paper cites.
Deng J, Dong W, Socher R, et al (2009) Imagenet: A large-scale hierarchical image database. In: CVPR
2009
Earlier work this paper cites.
Xiao J, Hays J, Ehinger KA, et al (2010) Sun database: Large-scale scene recognition from abbey to zoo. In: 2010 IEEE computer society conference on computer vision and pattern recognition, IEEE, pp 3485–3492
2010
Earlier work this paper cites.
Krizhevsky A, Sutskever I, Hinton GE (2012) Imagenet classification with deep convolutional neural networks. In: NIPS
2012
Earlier work this paper cites.
Parkhi OM, Vedaldi A, Zisserman A, et al (2012) Cats and dogs. In: 2012 IEEE conference on computer vision and pattern recognition, IEEE, pp 3498–3505
2012
Earlier work this paper cites.
Soomro K, Zamir AR, Shah M (2012) Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:12120402
2012
Earlier work this paper cites.
Krause J, Stark M, Deng J, et al (2013) 3d object representations for fine-grained categorization. In: Proceedings of the IEEE international conference on computer vision workshops, pp 554–561
2013
Earlier work this paper cites.
Maji S, Rahtu E, Kannala J, et al (2013) Fine-grained visual classification of aircraft. arXiv preprint arXiv:13065151
2013
Earlier work this paper cites.
Bossard L, Guillaumin M, Van Gool L (2014) Food-101–mining discriminative components with random forests. In: European conference on computer vision, Springer, pp 446–461
2014
Earlier work this paper cites.
Cimpoi M, Maji S, Kokkinos I, et al (2014) Describing textures in the wild. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp 3606–3613
2014
Earlier work this paper cites.
Long J, Shelhamer E, Darrell T (2015) Fully convolutional networks for semantic segmentation. In: CVPR
2015
Earlier work this paper cites.
Ren S, He K, Girshick R, et al (2015) Faster r-cnn: Towards real-time object detection with region proposal networks. In: NIPS
2015
Earlier work this paper cites.
Simonyan K, Zisserman A (2015) Very deep convolutional networks for large-scale image recognition. In: ICLR
2015
Earlier work this paper cites.
He K, Zhang X, Ren S, et al (2016) Deep residual learning for image recognition. In: CVPR
2016
Earlier work this paper cites.
Howard AG, Zhu M, Chen B, et al (2017) Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:170404861
2017
Earlier work this paper cites.
Vaswani A, Shazeer N, Parmar N, et al (2017) Attention is all you need. In: NIPS
2017
Earlier work this paper cites.
Anderson P, He X, Buehler C, et al (2018) Bottom-up and top-down attention for image captioning and visual question answering. In: CVPR
2018
Earlier work this paper cites.
Kim JH, Jun J, Zhang BT (2018) Bilinear attention networks. In: NIPS
2018
Earlier work this paper cites.
Devlin J, Chang MW, Lee K, et al (2019) Bert: Pre-training of deep bidirectional transformers for language understanding. In: NAACL-HLT
2019
Earlier work this paper cites.
Dong L, Yang N, Wang W, et al (2019) Unified language model pre-training for natural language understanding and generation. In: NeurIPS
2019
Earlier work this paper cites.
Gao P, Jiang Z, You H, et al (2019) Dynamic fusion with intra-and inter-modality attention flow for visual question answering. In: CVPR
2019
Earlier work this paper cites.
Helber P, Bischke B, Dengel A, et al (2019) Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 12(7):2217–2226
2019
Cited alongside, same era.
Houlsby N, Giurgiu A, Jastrzebski S, et al (2019) Parameter-efficient transfer learning for nlp. In: International Conference on Machine Learning, PMLR, pp 2790–2799
2019
Cited alongside, same era.
Lu J, Batra D, Parikh D, et al (2019) Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In: NeurIPS
2019
Cited alongside, same era.
Radford A, Wu J, Child R, et al (2019) Language models are unsupervised multitask learners. OpenAI blog
2019
Cited alongside, same era.
Recht B, Roelofs R, Schmidt L, et al (2019) Do imagenet classifiers generalize to imagenet? In: International Conference on Machine Learning, PMLR, pp 5389–5400
Touvron H, Cord M, Douze M, et al (2021) Training data-efficient image transformers & distillation through attention. In: ICML
2021
Closest in time.
Tsimpoukelli M, Menick JL, Cabi S, et al (2021) Multimodal few-shot learning with frozen language models. Advances in Neural Information Processing Systems 34:200–212
2021
Closest in time.
Yao Y, Zhang A, Zhang Z, et al (2021) Cpt: Colorful prompt tuning for pre-trained vision-language models. arXiv preprint arXiv:210911797
2021
Closest in time.
Alayrac JB, Donahue J, Luc P, et al (2022) Flamingo: a visual language model for few-shot learning. In: Oh AH, Agarwal A, Belgrave D, et al (eds) Advances in Neural Information Processing Systems
2022
Closest in time.
Gu Y, Han X, Liu Z, et al (2022) Ppt: Pre-trained prompt tuning for few-shot learning. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp 8410–8423
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
Tan H, Bansal M (2019) Lxmert: Learning cross-modality encoder representations from transformers. In: EMNLP-IJCNLP
2019
Cited alongside, same era.
Wang H, Ge S, Lipton Z, et al (2019) Learning robust global representations by penalizing local predictive power. Advances in Neural Information Processing Systems 32
2019
Cited alongside, same era.
Yu Z, Yu J, Cui Y, et al (2019) Deep modular co-attention networks for visual question answering. In: CVPR
2019
Cited alongside, same era.
Brown T, Mann B, Ryder N, et al (2020) Language models are few-shot learners. In: NeurIPS
2020
Cited alongside, same era.
Carion N, Massa F, Synnaeve G, et al (2020) End-to-end object detection with transformers. In: ECCV
2020
Cited alongside, same era.
Chen YC, Li L, Yu L, et al (2020) Uniter: Learning universal image-text representations. In: ECCV
2020
Cited alongside, same era.
Conneau A, Khandelwal K, Goyal N, et al (2020) Unsupervised cross-lingual representation learning at scale. In: ACL
2020
Cited alongside, same era.
2022
Closest in time.
He J, Zhou C, Ma X, et al (2022) Towards a unified view of parameter-efficient transfer learning. In: International Conference on Learning Representations
2022
Closest in time.
Hu S, Zhang Z, Ding N, et al (2022) Sparse structure search for parameter-efficient tuning. arXiv preprint arXiv:220607382
2022
Closest in time.
Jia M, Tang L, Chen BC, et al (2022) Visual prompt tuning. In: ECCV, pp 709–727
2022
Closest in time.
Li C, Liu H, Li LH, et al (2022) ELEVATER: A benchmark and toolkit for evaluating language-augmented visual models. In: Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track
2022
Closest in time.
Lian D, Zhou D, Feng J, et al (2022) Scaling & shifting your features: A new baseline for efficient model tuning. Advances in Neural Information Processing Systems 35:109–123
2022
Closest in time.
Lin Z, Geng S, Zhang R, et al (2022) Frozen clip models are efficient video learners. ECCV 2022
2022
Closest in time.
Sun T, Shao Y, Qian H, et al (2022) Black-box tuning for language-model-as-a-service. In: International Conference on Machine Learning, PMLR, pp 20,841–20,855
2022
Closest in time.
Wortsman M, Ilharco G, Kim JW, et al (2022) Robust fine-tuning of zero-shot models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 7959–7971
2022
Closest in time.
Yao Y, Chen Q, Zhang A, et al (2022) PEVL: Position-enhanced pre-training and prompt tuning for vision-language models. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp 11,104–11,117
2022
Closest in time.
Zhang R, Zeng Z, Guo Z (2022a) Can language understand depth? ACM MM 2022
2022
Closest in time.
Zhang R, Zhang W, Fang R, et al (2022b) Tip-adapter: Training-free adaption of clip for few-shot classification. In: ECCV 2022. Springer Nature Switzerland
2022
Closest in time.
Zhou K, Yang J, Loy CC, et al (2022) Learning to prompt for vision-language models. International Journal of Computer Vision pp 1–12
2022
Closest in time.
Liu P, Yuan W, Fu J, et al (2023) Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Computing Surveys 55(9):1–35
2023
Closest in time.
Zhang R, Hu X, Li B, et al (2023a) Prompt, generate, then cache: Cascade of foundation models makes strong few-shot learners. CVPR 2023
2023
Closest in time.
Zhang R, Wang L, Qiao Y, et al (2023b) Learning 3d representations from 2d pre-trained models via image-to-point masked autoencoders. CVPR 2023
2023
Closest in time.
Zhu X, Zhang R, He B, et al (2022) Pointclip v2: Adapting clip for powerful 3d open-world learning. ICCV 2023
2023
Closest in time.
Zhu X, Zhang R, He B, et al (2023) Not all features matter: Enhancing few-shot clip with adaptive prior refinement. ICCV 2023
2023
Closest in time.