Fetching the paper…
Reading the bibliography…
Language-image pre-training largely relies on how precisely and thoroughly a text describes its paired image.
Fei-Fei, L., Fergus, R., Perona, P.: Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories. In: IEEE Conf. Comput. Vis. Pattern Recog. (2004)
2004
Earlier work this paper cites.
Nilsback, M.E., Zisserman, A.: Automated flower classification over a large number of classes. In: Sixth Indian Conference on Computer Vision, Graphics & Image Processing (2008)
2008
Earlier work this paper cites.
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database. In: IEEE Conf. Comput. Vis. Pattern Recog. (2009)
2009
Earlier work this paper cites.
Krizhevsky, A., Hinton, G., et al.: Learning multiple layers of features from tiny images (2009)
2009
Earlier work this paper cites.
Xiao, J., Hays, J., Ehinger, K.A., Oliva, A., Torralba, A.: Sun database: Large-scale scene recognition from abbey to zoo. In: Int. Conf. Comput. Vis. (2010)
2010
Earlier work this paper cites.
Everingham, M., Winn, J.: The pascal visual object classes challenge 2012 (voc2012) development kit. Pattern Anal. Stat. Model. Comput. Learn., Tech. Rep 2007
2012
Earlier work this paper cites.
Parkhi, O.M., Vedaldi, A., Zisserman, A., Jawahar, C.: Cats and dogs. In: Int. Conf. Comput. Vis. (2012)
2012
Earlier work this paper cites.
Krause, J., Stark, M., Deng, J., Fei-Fei, L.: 3d object representations for fine-grained categorization. In: ICCVW (2013)
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
Bossard, L., Guillaumin, M., Van Gool, L.: Food-101–mining discriminative components with random forests. In: Eur. Conf. Comput. Vis. (2014)
2014
Earlier work this paper cites.
Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., Vedaldi, A.: Describing textures in the wild. In: IEEE Conf. Comput. Vis. Pattern Recog. (2014)
2014
Earlier work this paper cites.
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: Eur. Conf. Comput. Vis. pp. 740–755. Springer (2014)
2014
Earlier work this paper cites.
Mottaghi, R., Chen, X., Liu, X., Cho, N.G., Lee, S.W., Fidler, S., Urtasun, R., Yuille, A.: The role of context for object detection and semantic segmentation in the wild. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 891–898 (2014)
2014
Earlier work this paper cites.
Young, P., Lai, A., Hodosh, M., Hockenmaier, J.: From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics pp. 67–78 (2014)
2014
Earlier work this paper cites.
2016
Earlier work this paper cites.
Richter, S.R., Vineet, V., Roth, S., Koltun, V.: Playing for data: Ground truth from computer games. In: Eur. Conf. Comput. Vis. pp. 102–118. Springer (2016)
2016
Earlier work this paper cites.
Loshchilov, I., Hutter, F.: Decoupled weight decay regularization. arXiv:1711.05101 (2017)
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Caesar, H., Uijlings, J., Ferrari, V.: Coco-stuff: Thing and stuff classes in context. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 1209–1218 (2018)
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Sharma, P., Ding, N., Goodman, S., Soricut, R.: Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In: Association for Computational Linguistics (2018)
2018
Earlier work this paper cites.
Chen, Y., Li, W., Chen, X., Gool, L.V.: Learning semantic segmentation from synthetic data: A geometrically guided input-output adaptation approach. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 1841–1850 (2019)
2019
Earlier work this paper cites.
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., Rohrbach, M.: Towards vqa models that can read. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 8317–8326 (2019)
2019
Earlier work this paper cites.
Chen, T., Kornblith, S., Norouzi, M., Hinton, G.: A simple framework for contrastive learning of visual representations. In: Int. Conf. Mach. Learn. (2020)
2020
Cited alongside, same era.
Chefer, H., Gur, S., Wolf, L.: Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers. In: Int. Conf. Comput. Vis. pp. 397–406 (2021)
2021
Cited alongside, same era.
Chefer, H., Gur, S., Wolf, L.: Transformer interpretability beyond attention visualization. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 782–791 (2021)
2021
Cited alongside, same era.
Jia, C., Yang, Y., Xia, Y., Chen, Y.T., Parekh, Z., Pham, H., Le, Q., Sung, Y.H., Li, Z., Duerig, T.: Scaling up visual and vision-language representation learning with noisy text supervision. In: Int. Conf. Mach. Learn. (2021)
2021
Cited alongside, same era.
Dong, X., Bao, J., Zheng, Y., Zhang, T., Chen, D., Yang, H., Zeng, M., Zhang, W., Yuan, L., Chen, D., et al.: Maskclip: Masked self-distillation advances contrastive language-image pretraining. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 10995–11005 (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Geng, S., Yuan, J., Tian, Y., Chen, Y., Zhang, Y.: HiCLIP: Contrastive language-image pretraining with hierarchy-aware attention. In: Int. Conf. Learn. Represent. (2023)
2023
Later among the works it cites.
Kim, B., Jo, Y., Kim, J., Kim, S.: Misalign, contrast then distill: Rethinking misalignments in language-image pre-training. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 2563–2572 (2023)
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: Int. Conf. Mach. Learn. (2021)
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Dou, Z.Y., Kamath, A., Gan, Z., Zhang, P., Wang, J., Li, L., Liu, Z., Liu, C., LeCun, Y., Peng, N., et al.: Coarse-to-fine vision-language pre-training with fusion in the backbone. Adv. Neural Inform. Process. Syst. 35
2022
Cited alongside, same era.
Fürst, A., Rumetshofer, E., Lehner, J., Tran, V.T., Tang, F., Ramsauer, H., Kreil, D., Kopp, M., Klambauer, G., Bitto, A., et al.: Cloob: Modern hopfield networks with infoloob outperform clip. Adv. Neural Inform. Process. Syst. 35
2022
Cited alongside, same era.
Gao, Y., Liu, J., Xu, Z., Zhang, J., Li, K., Ji, R., Shen, C.: Pyramidclip: Hierarchical feature alignment for vision-language model pretraining. Adv. Neural Inform. Process. Syst. 35
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Li, J., Li, D., Xiong, C., Hoi, S.: Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In: Int. Conf. Mach. Learn. (2022)
2022
Cited alongside, same era.
2023
Later among the works it cites.
Li, Y., Fan, H., Hu, R., Feichtenhofer, C., He, K.: Scaling language-image pre-training via masking. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 23390–23400 (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Wu, S., Fei, H., Zhang, H., Chua, T.S.: Imagine that! abstract-to-intricate text-to-image synthesis with scene graph hallucination diffusion. Adv. Neural Inform. Process. Syst. 36
2023
Later among the works it cites.
Xu, M., Zhang, Z., Wei, F., Hu, H., Bai, X.: Side adapter network for open-vocabulary semantic segmentation. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 2945–2954 (2023)
2023
Later among the works it cites.
Yang, K., Deng, J., An, X., Li, J., Feng, Z., Guo, J., Yang, J., Liu, T.: Alip: Adaptive language-image pre-training with synthetic caption. In: Int. Conf. Comput. Vis. pp. 2922–2931 (2023)
2023
Later among the works it cites.
2023
Later among the works it cites.
Zhao, L., Zheng, K., Zheng, Y., Zhao, D., Zhou, J.: Rleg: Vision-language representation learning with diffusion-based embedding generation. Int. Conf. Mach. Learn. (2023)
2023
Later among the works it cites.
2024
Closest in time.
Dai, W., Li, J., Li, D., Tiong, A.M.H., Zhao, J., Wang, W., Li, B., Fung, P.N., Hoi, S.: InstructBLIP: Towards general-purpose vision-language models with instruction tuning. Adv. Neural Inform. Process. Syst. 36
2024
Closest in time.
Fan, L., Krishnan, D., Isola, P., Katabi, D., Tian, Y.: Improving clip training with language rewrites. Adv. Neural Inform. Process. Syst. 36
2024
Closest in time.
2024
Closest in time.
Hsieh, C.Y., Zhang, J., Ma, Z., Kembhavi, A., Krishna, R.: Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality. Adv. Neural Inform. Process. Syst. 36
2024
Closest in time.
Tian, Y., Fan, L., Isola, P., Chang, H., Krishnan, D.: Stablerep: Synthetic images from text-to-image models make strong visual representation learners. Adv. Neural Inform. Process. Syst. 36
2024
Closest in time.