Fetching the paper…
Reading the bibliography…
Pre-trained vision-language (V-L) models such as CLIP have shown excellent performance in many downstream cross-modal tasks.
Unsupervised cross-lingual representation learning at scale
Conneau, A.; Khandelwal, K.; Goyal, N.; Chaudhary, V.; Wenzek, G.; Guzmán, F.; Grave, E.; Ott, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1911
Earlier work this paper cites.
Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories
Fei-Fei, L.; Fergus, R.; and Perona, P. 2004 · 2004
Earlier work this paper cites.
Model compression
Buciluǎ, C.; Caruana, R.; and Niculescu-Mizil, A. 2006 · 2006
Earlier work this paper cites.
Automated flower classification over a large number of classes
Nilsback, M.-E.; and Zisserman, A. 2008 · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Cats and dogs
Parkhi, O. M.; Vedaldi, A.; Zisserman, A.; and Jawahar, C. 2012 · 2012
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Soomro, K.; Zamir, A. R.; and Shah, M. 2012 · 2012
Earlier work this paper cites.
Minilmv2: Multi-head self-attention relation distillation for compressing pretrained transformers
Wang, W.; Bao, H.; Huang, S.; Dong, L.; and Wei, F. 2020 · 2012
Earlier work this paper cites.
3d object representations for fine-grained categorization
Krause, J.; Stark, M.; Deng, J.; and Fei-Fei, L. 2013 · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft
Maji, S.; Rahtu, E.; Kannala, J.; Blaschko, M.; and Vedaldi, A. 2013 · 2013
Earlier work this paper cites.
Food-101–mining discriminative components with random forests
Bossard, L.; Guillaumin, M.; and Van Gool, L. 2014 · 2014
Earlier work this paper cites.
Fast r-cnn
Girshick, R. 2015 · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G.; Vinyals, O.; and Dean, J. 2015 · 2015
Earlier work this paper cites.
Training very deep networks
Srivastava, R. K.; Greff, K.; and Schmidhuber, J. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Cited alongside, same era.
Yfcc100m: The new data in multimedia research
Thomee, B.; Shamma, D. A.; Friedland, G.; Elizalde, B.; Ni, K.; Poland, D.; Borth, D.; and Li, L.-J. 2016 · 2016
Cited alongside, same era.
Knowledge distillation with feature maps for image classification
Chen, W.-C.; Chang, C.-C.; and Lee, C.-R. 2019 · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Oord, A. v. d.; Li, Y.; and Vinyals, O. 2018 · 2018
Cited alongside, same era.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
Helber, P.; Bischke, B.; Dengel, A.; and Borth, D. 2019 · 2019
Cited alongside, same era.
Astraea: Deploy ai services at the edge in elegant ways
Cross-lingual and multilingual clip
Carlsson, F.; Eisen, P.; Rekathati, F.; and Sahlgren, M. 2022 · 2022
Later among the works it cites.
Altclip: Altering the language encoder in clip for extended language capabilities
Chen, Z.; Liu, G.; Zhang, B.-W.; Ye, F.; Yang, Q.; and Wu, L. 2022 · 2022
Later among the works it cites.
Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark
Gu, J.; Meng, X.; Lu, G.; Hou, L.; Minzhe, N.; Liang, X.; Yao, L.; Huang, R.; Zhang, W.; Jiang, X.; et al. 2022 · 2022
Later among the works it cites.
TinyML Techniques for running Machine Learning models on Edge Devices
Mukherjee, A.; Ukil, A.; Dey, S.; and Kulkarni, G. 2022 · 2022
Later among the works it cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C.; Beaumont, R.; Vencu, R.; Gordon, C.; Wightman, R.; Cherti, M.; Coombes, T.; Katta, A.; Mullis, C.; Wortsman, M.; et al. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fu, Z.; Yang, J.; Bai, C.; Chen, X.; Zhang, C.; Zhang, Y.; and Wang, D. 2020 · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. 2020 · 2020
Cited alongside, same era.
Contrastive language-image pre-training for the italian language
Bianchi, F.; Attanasio, G.; Pisoni, R.; Terragni, S.; Sarti, G.; and Lakshmi, S. 2021 · 2021
Cited alongside, same era.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Changpinyo, S.; Sharma, P.; Ding, N.; and Soricut, R. 2021 · 2021
Cited alongside, same era.
Pre-Training with Whole Word Masking for Chinese BERT
Cui, Y.; Che, W.; Liu, T.; Qin, B.; and Yang, Z. 2021 · 2021
Cited alongside, same era.
Knowledge distillation: A survey
Gou, J.; Yu, B.; Maybank, S. J.; and Tao, D. 2021 · 2021
Cited alongside, same era.
Open-vocabulary object detection via vision and language knowledge distillation
Gu, X.; Lin, T.-Y.; Kuo, W.; and Cui, Y. 2021 · 2021
Cited alongside, same era.
Flava: A foundational language and vision alignment model
Singh, A.; Hu, R.; Goswami, V.; Couairon, G.; Galuba, W.; Rohrbach, M.; and Kiela, D. 2022 · 2022
Later among the works it cites.
Contrastive learning rivals masked image modeling in fine-tuning via feature distillation
Wei, Y.; Hu, H.; Xie, Z.; Zhang, Z.; Cao, Y.; Bao, J.; Chen, D.; and Guo, B. 2022 · 2022
Later among the works it cites.
Groupvit: Semantic segmentation emerges from text supervision
Xu, J.; De Mello, S.; Liu, S.; Byeon, W.; Breuel, T.; Kautz, J.; and Wang, X. 2022 · 2022
Later among the works it cites.
Lit: Zero-shot transfer with locked-image text tuning
Zhai, X.; Wang, X.; Mustafa, B.; Steiner, A.; Keysers, D.; Kolesnikov, A.; and Beyer, L. 2022 · 2022
Later among the works it cites.
Learning to prompt for vision-language models
Zhou, K.; Yang, J.; Loy, C. C.; and Liu, Z. 2022 · 2022
Later among the works it cites.
Maple: Multi-modal prompt learning
Khattak, M. U.; Rasheed, H.; Maaz, M.; Khan, S.; and Khan, F. S. 2023 · 2023
Later among the works it cites.
Tiny Machine Learning: Progress and Futures [Feature]
Lin, J.; Zhu, L.; Chen, W.-M.; Wang, W.-C.; and Han, S. 2023 · 2023
Later among the works it cites.
Tinyclip: Clip distillation via affinity mimicking and weight inheritance
Wu, K.; Peng, H.; Zhou, Z.; Xiao, B.; Liu, M.; Yuan, L.; Xuan, H.; Valenzuela, M.; Chen, X. S.; Wang, X.; et al. 2023 · 2023
Later among the works it cites.
A survey on knowledge distillation of large language models
Xu, X.; Li, M.; Tao, C.; Shen, T.; Cheng, R.; Li, J.; Xu, C.; Tao, D.; and Zhou, T. 2024 · 2024
Closest in time.