Fetching the paper…
Reading the bibliography…
Contrastive vision-language models like CLIP have shown great progress in transfer learning.
Billion-scale semi-supervised learning for image classification
Yalniz, I. Z.; J’egou, H.; Chen, K.; Paluri, M.; and Mahajan, D. 2019 · 1905
Earlier work this paper cites.
Revisiting self-training for neural sequence generation
He, J.; Gu, J.; Shen, J.; and Ranzato, M. 2019 · 1909
Earlier work this paper cites.
Probability of error of some adaptive pattern-recognition machines
Scudder, H. 1965 · 1965
Earlier work this paper cites.
Unsupervised word sense disambiguation rivaling supervised methods
Yarowsky, D. 1995 · 1995
Earlier work this paper cites.
Automatically generating extraction patterns from untagged text
Riloff, E. 1996 · 1996
Earlier work this paper cites.
Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories
Fei-Fei, L.; Fergus, R.; and Perona, P. 2004 · 2004
Earlier work this paper cites.
A simple semi-supervised learning framework for object detection
Sohn, K.; Zhang, Z.; Li, C.-L.; Zhang, H.; Lee, C.-Y.; and Pfister, T. 2020 · 2005
Earlier work this paper cites.
Automated flower classification over a large number of classes
2008 · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Shin, T.; Razeghi, Y.; Logan IV, R. L.; Wallace, E.; and Singh, S. 2020 · 2010
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo
Xiao, J.; Hays, J.; Ehinger, K. A.; Oliva, A.; and Torralba, A. 2010 · 2010
Earlier work this paper cites.
Cats and dogs
Parkhi, O. M.; Vedaldi, A.; Zisserman, A.; and Jawahar, C. 2012 · 2012
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Soomro, K.; Zamir, A. R.; and Shah, M. 2012 · 2012
Earlier work this paper cites.
3d object representations for fine-grained categorization
Krause, J.; Stark, M.; Deng, J.; and Fei-Fei, L. 2013 · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft
Maji, S.; Rahtu, E.; Kannala, J.; Blaschko, M.; and Vedaldi, A. 2013 · 2013
Earlier work this paper cites.
Food-101–mining discriminative components with random forests
Bossard, L.; Guillaumin, M.; and Van Gool, L. 2014 · 2014
Earlier work this paper cites.
Describing textures in the wild
Cimpoi, M.; Maji, S.; Kokkinos, I.; Mohamed, S.; and Vedaldi, A. 2014 · 2014
Cited alongside, same era.
Neural machine translation of rare words with subword units
Sennrich, R.; Haddow, B.; and Birch, A. 2015 · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Representation learning with contrastive predictive coding
Van den Oord, A.; Li, Y.; and Vinyals, O. 2018 · 2018
Cited alongside, same era.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
Helber, P.; Bischke, B.; Dengel, A.; and Borth, D. 2019 · 2019
Scaling up visual and vision-language representation learning with noisy text supervision
Jia, C.; Yang, Y.; Xia, Y.; Chen, Y.-T.; Parekh, Z.; Pham, H.; Le, Q.; Sung, Y.-H.; Li, Z.; and Duerig, T. 2021 · 2021
Later among the works it cites.
The power of scale for parameter-efficient prompt tuning
Lester, B.; Al-Rfou, R.; and Constant, N. 2021 · 2021
Later among the works it cites.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L.; and Liang, P. 2021 · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 · 2021
Later among the works it cites.
Clip4clip: An empirical study of clip for end to end video clip retrieval
Luo, H.; Ji, L.; Zhong, M.; Chen, Y.; Lei, W.; Duan, N.; and Li, T. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Lessons from building acoustic models with a million hours of speech
Parthasarathi, S. H. K.; and Strom, N. 2019 · 2019
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
Chen, T.; Kornblith, S.; Norouzi, M.; and Hinton, G. 2020 · 2020
Cited alongside, same era.
Bootstrap your own latent-a new approach to self-supervised learning
Grill, J.-B.; Strub, F.; Altché, F.; Tallec, C.; Richemond, P.; Buchatskaya, E.; Doersch, C.; Avila Pires, B.; Guo, Z.; Gheshlaghi Azar, M.; et al. 2020 · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. 2020 · 2020
Cited alongside, same era.
How can we know what language models know?
Jiang, Z.; Xu, F. F.; Araki, J.; and Neubig, G. 2020 · 2020
Cited alongside, same era.
Self-training for end-to-end speech recognition
Kahn, J.; Lee, A.; and Hannun, A. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Later among the works it cites.
Rizve, M. N.; Duarte, K.; Rawat, Y. S.; and Shah, M. 2021 · 2021
Later among the works it cites.
CLIP4Caption: CLIP for Video Caption
Tang, M.; Wang, Z.; Liu, Z.; Rao, F.; Li, D.; and Li, X. 2021 · 2021
Later among the works it cites.
Actionclip: A new paradigm for video action recognition
Wang, M.; Xing, J.; and Liu, Y. 2021 · 2021
Later among the works it cites.
FILIP: Fine-grained Interactive Language-Image Pre-Training
Yao, L.; Huang, R.; Hou, L.; Lu, G.; Niu, M.; Xu, H.; Liang, X.; Li, Z.; Jiang, X.; and Xu, C. 2021 · 2021
Later among the works it cites.
Florence: A New Foundation Model for Computer Vision
Yuan, L.; Chen, D.; Chen, Y.-L.; Codella, N.; Dai, X.; Gao, J.; Hu, H.; Huang, X.; Li, B.; Li, C.; et al. 2021 · 2021
Later among the works it cites.
Factual probing is [mask]: Learning vs. learning to recall
Zhong, Z.; Friedman, D.; and Chen, D. 2021 · 2021
Later among the works it cites.
Learning to prompt for vision-language models
Zhou, K.; Yang, J.; Loy, C. C.; and Liu, Z. 2021 · 2021
Later among the works it cites.
Conditional prompt learning for vision-language models
Zhou, K.; Yang, J.; Loy, C. C.; and Liu, Z. 2022 · 2021
Later among the works it cites.
Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language Model
Du, Y.; Wei, F.; Zhang, Z.; Shi, M.; Gao, Y.; and Li, G. 2022 · 2022
Closest in time.
Wukong: 100 Million Large-scale Chinese Cross-modal Pre-training Dataset and A Foundation Framework
Gu, J.; Meng, X.; Lu, G.; Hou, L.; Niu, M.; Xu, H.; Liang, X.; Zhang, W.; Jiang, X.; and Xu, C. 2022 · 2022
Closest in time.