Fetching the paper…
Reading the bibliography…
Vision-language pre-training like CLIP has shown promising performance on various downstream tasks such as zero-shot image classification and image-text retrieval.
Improved baselines with momentum contrastive learning
Chen, X., Fan, H., Girshick, R., and He, K · 2003
Earlier work this paper cites.
Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories
Fei-Fei, L., Fergus, R., and Perona, P · 2004
Earlier work this paper cites.
Automated flower classification over a large number of classes
Nilsback, M.-E. and Zisserman, A · 2008
Earlier work this paper cites.
Visualizing data using t-sne
Van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo
Xiao, J., Hays, J., Ehinger, K. A., Oliva, A., and Torralba, A · 2010
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning
Coates, A., Ng, A., and Lee, H · 2011
Earlier work this paper cites.
Unimo: Towards unified-modal understanding and generation via cross-modal contrastive learning
Li, W., Gao, C., Niu, G., Xiao, X., Liu, H., Liu, J., Wu, H., and Wang, H · 2012
Earlier work this paper cites.
Cats and dogs
Parkhi, O. M., Vedaldi, A., Zisserman, A., and Jawahar, C · 2012
Earlier work this paper cites.
Challenges in representation learning: A report on three machine learning contests
Goodfellow, I. J., Erhan, D., Carrier, P. L., Courville, A., Mirza, M., Hamner, B., Cukierski, W., Tang, Y., Thaler, D., Lee, D.-H., et al · 2013
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
Hodosh, M., Young, P., and Hockenmaier, J · 2013
Earlier work this paper cites.
Food-101–mining discriminative components with random forests
Bossard, L., Guillaumin, M., and Van Gool, L · 2014
Earlier work this paper cites.
Describing textures in the wild
Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., and Vedaldi, A · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Fast r-cnn
Girshick, R · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
End-to-end people detection in crowded scenes
Stewart, R., Andriluka, M., and Ng, A. Y · 2016
Earlier work this paper cites.
Yfcc100m: The new data in multimedia research
Thomee, B., Shamma, D. A., Friedland, G., Elizalde, B., Ni, K., Poland, D., Borth, D., and Li, L.-J · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Cited alongside, same era.
Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., Ginsburg, B., Houston, M., Kuchaiev, O., Venkatesh, G., et al · 2017
Cited alongside, same era.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Vilt: Vision-and-language transformer without convolution or region supervision
Kim, W., Son, B., and Kim, I · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Later among the works it cites.
Filip: fine-grained interactive language-image pre-training
Yao, L., Huang, R., Hou, L., Lu, G., Niu, M., Xu, H., Liang, X., Li, Z., Jiang, X., and Xu, C · 2021
Later among the works it cites.
Vinvl: Revisiting visual representations in vision-language models
Zhang, P., Li, X., Hu, X., Yang, J., Zhang, L., Wang, L., Choi, Y., and Gao, J · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Cited alongside, same era.
Mobilenetv2: Inverted residuals and linear bottlenecks
Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Cited alongside, same era.
Visualbert: A simple and performant baseline for vision and language
Li, L. H., Yatskar, M., Yin, D., Hsieh, C.-J., and Chang, K.-W · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J., Batra, D., Parikh, D., and Lee, S · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Cited alongside, same era.
End-to-end object detection with transformers
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S · 2020
Cited alongside, same era.
Robust cross-modal representation learning with progressive self-distillation
Andonian, A., Chen, S., and Hamid, R · 2022
Later among the works it cites.
Cui, Y., Zhao, L., Liang, F., Li, Y., and Shao, J · 2022
Later among the works it cites.
An empirical study of training end-to-end vision-and-language transformers
Dou, Z.-Y., Xu, Y., Gan, Z., Wang, J., Wang, S., Wang, L., Zhu, C., Zhang, P., Yuan, L., Peng, N., et al · 2022
Later among the works it cites.
Pyramidclip: Hierarchical feature alignment for vision-language model pretraining
Gao, Y., Liu, J., Xu, Z., Zhang, J., Li, K., Ji, R., and Shen, C · 2022
Later among the works it cites.
Cyclip: Cyclic contrastive language-image pretraining
Goel, S., Bansal, H., Bhatia, S., Rossi, R., Vinay, V., and Grover, A · 2022
Later among the works it cites.
Nlip: Noise-robust language-image pre-training
Huang, R., Long, Y., Han, J., Xu, H., Liang, X., Xu, C., and Liang, X · 2022
Later among the works it cites.
Uniclip: Unified framework for contrastive language-image pre-training
Lee, J., Kim, J., Shon, H., Kim, B., Kim, S. H., Lee, H., and Kim, J · 2022
Later among the works it cites.
Cots: Collaborative two-stream vision-language pre-training model for cross-modal retrieval
Lu, H., Fei, N., Huo, Y., Gao, Y., Lu, Z., and Wen, J.-R · 2022
Later among the works it cites.
Slip: Self-supervision meets language-image pre-training
Mu, N., Kirillov, A., Wagner, D., and Xie, S · 2022
Later among the works it cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al · 2022
Later among the works it cites.
Coca: Contrastive captioners are image-text foundation models
Yu, J., Wang, Z., Vasudevan, V., Yeung, L., Seyedhosseini, M., and Wu, Y · 2022
Later among the works it cites.
Tokenflow: Rethinking fine-grained cross-modal alignment in vision-language retrieval
Zou, X., Wu, C., Cheng, L., and Wang, Z · 2022
Later among the works it cites.
Softclip: Softer cross-modal alignment makes clip stronger
Gao, Y., Liu, J., Xu, Z., Wu, T., Liu, W., Yang, J., Li, K., and Sun, X · 2023
Closest in time.
Filtering, distillation, and hard negatives for vision-language pre-training
Radenovic, F., Dubey, A., Kadian, A., Mihaylov, T., Vandenhende, S., Patel, Y., Wen, Y., Ramanathan, V., and Mahajan, D · 2023
Closest in time.