Fetching the paper…
Reading the bibliography…
Large pre-trained vision-language models like CLIP have shown great potential in learning representations that are transferable across a wide range of downstream tasks.
Fei-Fei L, Fergus R, Perona P (2004) Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories. In: CVPR-W
2004
Earlier work this paper cites.
Nilsback ME, Zisserman A (2008) Automated flower classification over a large number of classes. In: ICVGIP
2008
Earlier work this paper cites.
Deng J, Dong W, Socher R, Li LJ, Li K, Fei-Fei L (2009) Imagenet: A large-scale hierarchical image database. In: CVPR
2009
Earlier work this paper cites.
Xiao J, Hays J, Ehinger KA, Oliva A, Torralba A (2010) Sun database: Large-scale scene recognition from abbey to zoo. In: CVPR
2010
Earlier work this paper cites.
Parkhi OM, Vedaldi A, Zisserman A, Jawahar C (2012) Cats and dogs. In: CVPR
2012
Earlier work this paper cites.
Soomro K, Zamir AR, Shah M (2012) Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:12120402
2012
Earlier work this paper cites.
Elhoseiny M, Saleh B, Elgammal A (2013) Write a classifier: Zero-shot learning using purely textual descriptions. In: ICCV
2013
Earlier work this paper cites.
Frome A, Corrado G, Shlens J, Bengio S, Dean J, Ranzato M, Mikolov T (2013) Devise: A deep visual-semantic embedding model. In: NeurIPS
2013
Earlier work this paper cites.
Krause J, Stark M, Deng J, Fei-Fei L (2013) 3d object representations for fine-grained categorization. In: ICCV-W
2013
Earlier work this paper cites.
Maji S, Rahtu E, Kannala J, Blaschko M, Vedaldi A (2013) Fine-grained visual classification of aircraft. arXiv preprint arXiv:13065151
2013
Earlier work this paper cites.
Socher R, Ganjoo M, Sridhar H, Bastani O, Manning CD, Ng AY (2013) Zero-shot learning through cross-modal transfer. In: NeurIPS
2013
Earlier work this paper cites.
Bossard L, Guillaumin M, Van Gool L (2014) Food-101–mining discriminative components with random forests. In: ECCV
2014
Earlier work this paper cites.
Cimpoi M, Maji S, Kokkinos I, Mohamed S, Vedaldi A (2014) Describing textures in the wild. In: CVPR
2014
Earlier work this paper cites.
Lei Ba J, Swersky K, Fidler S, et al. (2015) Predicting deep zero-shot convolutional neural networks using textual descriptions. In: ICCV
2015
Earlier work this paper cites.
He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: CVPR
2016
Earlier work this paper cites.
Joulin A, Van Der Maaten L, Jabri A, Vasilache N (2016) Learning visual features from large weakly supervised data. In: ECCV
2016
Earlier work this paper cites.
Sennrich R, Haddow B, Birch A (2016) Neural machine translation of rare words with subword units. In: ACL
2016
Earlier work this paper cites.
Gomez L, Patel Y, Rusiñol M, Karatzas D, Jawahar C (2017) Self-supervised learning of visual features through embedding images into text topic spaces. In: CVPR
2017
Earlier work this paper cites.
Li A, Jabri A, Joulin A, van der Maaten L (2017) Learning visual n-grams from web data. In: ICCV
2017
Cited alongside, same era.
Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I (2017) Attention is all you need. In: NeurIPS
2017
Cited alongside, same era.
Zhou B, Lapedriza A, Khosla A, Oliva A, Torralba A (2017) Places: A 10 million image database for scene recognition. IEEE transactions on pattern analysis and machine intelligence 40(6):1452–1464
2017
Cited alongside, same era.
Helber P, Bischke B, Dengel A, Borth D (2019) Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing
2019
Cited alongside, same era.
Petroni F, Rocktäschel T, Lewis P, Bakhtin A, Wu Y, Miller AH, Riedel S (2019) Language models as knowledge bases? In: EMNLP
Zhang Y, Jiang H, Miura Y, Manning CD, Langlotz CP (2020) Contrastive learning of medical visual representations from paired images and text. arXiv preprint arXiv:201000747
2020
Later among the works it cites.
Bommasani R, Hudson DA, Adeli E, Altman R, Arora S, von Arx S, Bernstein MS, Bohg J, Bosselut A, Brunskill E, et al. (2021) On the opportunities and risks of foundation models. arXiv preprint arXiv:210807258
2021
Closest in time.
Desai K, Johnson J (2021) Virtex: Learning visual representations from textual annotations. In: CVPR
2021
Closest in time.
Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, Unterthiner T, Dehghani M, Minderer M, Heigold G, Gelly S, et al. (2021) An image is worth 16x16 words: Transformers for image recognition at scale. In: ICLR
2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
Recht B, Roelofs R, Schmidt L, Shankar V (2019) Do imagenet classifiers generalize to imagenet? In: ICML
2019
Cited alongside, same era.
Wang H, Ge S, Lipton Z, Xing EP (2019) Learning robust global representations by penalizing local predictive power. In: NeurIPS
2019
Cited alongside, same era.
Brown TB, Mann B, Ryder N, Subbiah M, Kaplan J, Dhariwal P, Neelakantan A, Shyam P, Sastry G, Askell A, et al. (2020) Language models are few-shot learners. arXiv preprint arXiv:200514165
2020
Cited alongside, same era.
Chen T, Kornblith S, Norouzi M, Hinton G (2020) A simple framework for contrastive learning of visual representations. In: ICML
2020
Cited alongside, same era.
Gao T, Fisch A, Chen D (2020) Making pre-trained language models better few-shot learners. arXiv preprint arXiv:201215723
2020
Cited alongside, same era.
He K, Fan H, Wu Y, Xie S, Girshick R (2020) Momentum contrast for unsupervised visual representation learning. In: CVPR
2020
Cited alongside, same era.
Hénaff OJ, Srinivas A, Fauw JD, Razavi A, Doersch C, Eslami SMA, van den Oord A (2020) Data-efficient image recognition with contrastive predictive coding. In: ICML
2020
Cited alongside, same era.
Fürst A, Rumetshofer E, Tran V, Ramsauer H, Tang F, Lehner J, Kreil D, Kopp M, Klambauer G, Bitto-Nemling A, et al. (2021) Cloob: Modern hopfield networks with infoloob outperform clip. arXiv preprint arXiv:211011316
2021
Closest in time.
Gao P, Geng S, Zhang R, Ma T, Fang R, Zhang Y, Li H, Qiao Y (2021) Clip-adapter: Better vision-language models with feature adapters. arXiv preprint arXiv:211004544
2021
Closest in time.
Jia C, Yang Y, Xia Y, Chen YT, Parekh Z, Pham H, Le QV, Sung Y, Li Z, Duerig T (2021) Scaling up visual and vision-language representation learning with noisy text supervision. In: ICML
2021
Closest in time.
Lester B, Al-Rfou R, Constant N (2021) The power of scale for parameter-efficient prompt tuning. arXiv preprint arXiv:210408691
2021
Closest in time.
Li XL, Liang P (2021) Prefix-tuning: Optimizing continuous prompts for generation. arXiv preprint arXiv:210100190
2021
Closest in time.
Li Y, Liang F, Zhao L, Cui Y, Ouyang W, Shao J, Yu F, Yan J (2021) Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm. arXiv preprint arXiv:211005208
2021
Closest in time.
Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, et al. (2021) Learning transferable visual models from natural language supervision. In: ICML
2021
Closest in time.
Singh A, Hu R, Goswami V, Couairon G, Galuba W, Rohrbach M, Kiela D (2021) Flava: A foundational language and vision alignment model. arXiv preprint arXiv:211204482
2021
Closest in time.
Yuan L, Chen D, Chen YL, Codella N, Dai X, Gao J, Hu H, Huang X, Li B, Li C, et al. (2021) Florence: A new foundation model for computer vision. arXiv preprint arXiv:211111432
2021
Closest in time.
Zhong Z, Friedman D, Chen D (2021) Factual probing is [mask]: Learning vs. learning to recall. In: NAACL
2021
Closest in time.
Zhou K, Liu Z, Qiao Y, Xiang T, Loy CC (2021) Domain generalization: A survey. arXiv preprint arXiv:210302503
2021
Closest in time.
Bahng H, Jahanian A, Sankaranarayanan S, Isola P (2022) Visual prompting: Modifying pixel space to adapt pre-trained models. arXiv preprint arXiv:220317274
2022
Closest in time.
Jia M, Tang L, Chen BC, Cardie C, Belongie S, Hariharan B, Lim SN (2022) Visual prompt tuning. arXiv preprint arXiv:220312119
2022
Closest in time.