Fetching the paper…
Reading the bibliography…
Contrastive Language-Image Pre-training (CLIP) has drawn increasing attention recently for its transferable visual representation learning.
“Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories,”
FF Li, R. Fergus, and P. Perona, · 2004
Earlier work this paper cites.
“Cats and dogs,”
Andrea Vedaldi, · 2012
Earlier work this paper cites.
“Describing textures in the wild,”
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi, · 2013
Earlier work this paper cites.
“Food-101 – mining discriminative components with random forests,”
L. Bossard, M. Guillaumin, and L. V. Gool, · 2014
Earlier work this paper cites.
“Deep residual learning for image recognition,”
K. He, X. Zhang, S. Ren, and J. Sun, · 2016
Earlier work this paper cites.
“Mask r-cnn,”
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming Wei Chang, Kenton Lee, and Kristina Toutanova, · 2018
Cited alongside, same era.
“Parameter-efficient transfer learning for nlp,”
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly, · 2019
Cited alongside, same era.
“Uniter: Learning universal image-text representations,”
Y. C. Chen, L. Li, L. Yu, A. E. Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu, · 2019
Cited alongside, same era.
“How can we know what language models know?,”
Z. Jiang, F. F. Xu, J. Araki, and G. Neubig, · 2019
Cited alongside, same era.
“End-to-end object detection with transformers,”
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, · 2020
Cited alongside, same era.
“Autoprompt: Eliciting knowledge from language models with automatically generated prompts,”
“An image is worth 16x16 words: Transformers for image recognition at scale,”
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, and N. Houlsby, · 2020
Later among the works it cites.
“Learning transferable visual models from natural language supervision,”
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al., · 2021
Closest in time.
“Simvlm: Simple visual language model pretraining with weak supervision,”
Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, and Yuan Cao, · 2021
Closest in time.
“Scaling up visual and vision-language representation learning with noisy text supervision,”
C. Jia, Y. Yang, Y. Xia, Y. T. Chen, and T. Duerig, · 2021
Closest in time.
“Prefix-tuning: Optimizing continuous prompts for generation,”
X. L. Li and P. Liang, · 2021
Closest in time.
“The power of scale for parameter-efficient prompt tuning,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Shin, Y. Razeghi, Irl Logan, E. Wallace, and S. Singh, · 2020
Cited alongside, same era.
“Oscar: Object-semantics aligned pre-training for vision-language tasks,”
X. Li, X. Yin, C. Li, P. Zhang, X. Hu, L. Zhang, L. Wang, H. Hu, L. Dong, and F. Wei, · 2020
Cited alongside, same era.
B. Lester, R. Al-Rfou, and N. Constant, · 2021
Closest in time.
“Learning to prompt for vision-language models,”
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu, · 2022
Closest in time.