Fetching the paper…
Reading the bibliography…
We present SynthCLIP, a CLIP model trained on entirely synthetic text-image pairs.
Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories
Fei-Fei, L., Fergus, R., and Perona, P. (2004) · 2004
Earlier work this paper cites.
Automated flower classification over a large number of classes
Nilsback, M.-E. and Zisserman, A. (2008) · 2008
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo
Xiao, J., Hays, J., Ehinger, K. A., Oliva, A., and Torralba, A. (2010) · 2010
Earlier work this paper cites.
Cats and dogs
Parkhi, O. M., Vedaldi, A., Zisserman, A., and Jawahar, C. (2012) · 2012
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
Hodosh, M., Young, P., and Hockenmaier, J. (2013) · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft
Maji, S., Rahtu, E., Kannala, J., Blaschko, M., and Vedaldi, A. (2013) · 2013
Earlier work this paper cites.
Food-101 – mining discriminative components with random forests
Bossard, L., Guillaumin, M., and Van Gool, L. (2014) · 2014
Earlier work this paper cites.
Describing textures in the wild
Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., , and Vedaldi, A. (2014) · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L. (2014) · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P., Lai, A., Hodosh, M., and Hockenmaier, J. (2014) · 2014
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L. (2015) · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
Schroff, F., Kalenichenko, D., and Philbin, J. (2015) · 2015
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
Noroozi, M. and Favaro, P. (2016) · 2016
Earlier work this paper cites.
Context encoders: Feature learning by inpainting
Pathak, D., Krahenbuhl, P., Donahue, J., Darrell, T., and Efros, A. A. (2016) · 2016
Earlier work this paper cites.
Playing for data: Ground truth from computer games
Richter, S. R., Vineet, V., Roth, S., and Koltun, V. (2016) · 2016
Earlier work this paper cites.
The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes
Ros, G., Sellart, L., Materzynska, J., Vazquez, D., and Lopez, A. M. (2016) · 2016
Earlier work this paper cites.
Driving in the matrix: Can virtual worlds replace human-generated annotations for real world tasks?
Johnson-Roberson, M., Barto, C., Mehta, R., Sridhar, S. N., Rosaen, K., and Vasudevan, R. (2017) · 2017
Earlier work this paper cites.
Learning from synthetic humans
Varol, G., Romero, J., Martin, X., Mahmood, N., Black, M. J., Laptev, I., and Schmid, C. (2017) · 2017
Earlier work this paper cites.
Unsupervised representation learning by predicting image rotations
Gidaris, S., Singh, P., and Komodakis, N. (2018) · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R. (2018) · 2018
Earlier work this paper cites.
How good is my gan?
Shmelkov, K., Schmid, C., and Alahari, K. (2018) · 2018
Earlier work this paper cites.
Learning semantic segmentation from synthetic data: A geometrically guided input-output adaptation approach
Chen, Y., Li, W., Chen, X., and Van Gool, L. (2019) · 2019
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. (2020) · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. (2020) · 2020
Cited alongside, same era.
Generating synthetic audio data for attention-based speech recognition systems
Rossenbach, N., Zeyer, A., Schlüter, R., and Ney, H. (2020) · 2020
Cited alongside, same era.
Generative data augmentation for commonsense reasoning
Yang, Y., Malaviya, C., Fernandez, J., Swayamdipta, S., Le Bras, R., Wang, J.-P., Bhagavatula, C., Choi, Y., and Downey, D. (2020) · 2020
Cited alongside, same era.
Don’t generate me: Training differentially private generative models with sinkhorn divergence
Cao, T., Bie, A., Vahdat, A., Fidler, S., and Kreis, K. (2021) · 2021
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A. (2021) · 2021
Cited alongside, same era.
Improved baselines for vision-language pre-training
Fini, E., Astolfi, P., Romero-Soriano, A., Verbeek, J., and Drozdzal, M. (2023) · 2023
Later among the works it cites.
Llm blueprint: Enabling text-to-image generation with complex and detailed prompts
Gani, H., Bhat, S. F., Naseer, M., Khan, S., and Wonka, P. (2023) · 2023
Later among the works it cites.
A systematic survey of prompt engineering on vision-language foundation models
Gu, J., Han, Z., Chen, S., Beirami, A., He, B., Zhang, G., Liao, R., Qin, Y., Tresp, V., and Torr, P. (2023) · 2023
Later among the works it cites.
Is synthetic data from generative models ready for image recognition?
He, R., Sun, S., Yu, X., Xue, C., Zhang, W., Torr, P., Bai, S., and QI, X. (2023) · 2023
Later among the works it cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al. (2023) · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Changpinyo, S., Sharma, P., Ding, N., and Soricut, R. (2021) · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. (2021) · 2021
Cited alongside, same era.
The web is your oyster-knowledge-intensive nlp against a very large web corpus
Piktus, A., Petroni, F., Karpukhin, V., Okhonko, D., Broscheit, S., Izacard, G., Lewis, P., Oğuz, B., Grave, E., Yih, W.-t., et al. (2021) · 2021
Cited alongside, same era.
Multi-task learning for dense prediction tasks: A survey
Vandenhende, S., Georgoulis, S., Van Gansbeke, W., Proesmans, M., Dai, D., and Van Gool, L. (2021) · 2021
Cited alongside, same era.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R. (2022) · 2022
Cited alongside, same era.
Generative models as a data source for multiview representation learning
Jahanian, A., Puig, X., Tian, Y., and Isola, P. (2022) · 2022
Cited alongside, same era.
Clipstyler: Image style transfer with a single text condition
Kwon, G. and Ye, J. C. (2022) · 2022
Cited alongside, same era.
Later among the works it cites.
Noise-aware learning from web-crawled image-text data for image captioning
Kang, W., Mun, J., Lee, S., and Roh, B. (2023) · 2023
Later among the works it cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I. (2023) · 2023
Later among the works it cites.
From scarcity to efficiency: Improving clip training via visual-enriched captions
Lai, Z., Zhang, H., Wu, W., Bai, H., Timofeev, A., Du, X., Gan, Z., Shan, J., Chuah, C.-N., Yang, Y., et al. (2023) · 2023
Later among the works it cites.
Fake it till you make it: Learning transferable representations from synthetic imagenet clones
Sariyildiz, M. B., Alahari, K., Larlus, D., and Kalantidis, Y. (2023) · 2023
Later among the works it cites.
Adversarial diffusion distillation
Sauer, A., Lorenz, D., Blattmann, A., and Rombach, R. (2023) · 2023
Later among the works it cites.
Safe latent diffusion: Mitigating inappropriate degeneration in diffusion models
Schramowski, P., Brack, M., Deiseroth, B., and Kersting, K. (2023) · 2023
Later among the works it cites.
Stablerep: Synthetic images from text-to-image models make strong visual representation learners
Tian, Y., Fan, L., Isola, P., Chang, H., and Krishnan, D. (2023) · 2023
Later among the works it cites.
Paragraph-to-image generation with information-enriched diffusion model
Wu, W., Li, Z., He, Y., Shou, M. Z., Shen, C., Cheng, L., Li, Y., Gao, T., Zhang, D., and Wang, Z. (2023) · 2023
Later among the works it cites.
Xu, H., Xie, S., Tan, X. E., Huang, P.-Y., Howes, R., Sharma, V., Li, S.-W., Ghosh, G., Zettlemoyer, L., and Feichtenhofer, C. (2023) · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training
Zhai, X., Mustafa, B., Kolesnikov, A., and Beyer, L. (2023) · 2023
Later among the works it cites.
Training on thin air: Improve image classification with generated data
Zhou, Y., Sahak, H., and Ba, J. (2023) · 2023
Later among the works it cites.
Scaling laws of synthetic images for model training… for now
Fan, L., Chen, K., Krishnan, D., Katabi, D., Isola, P., and Tian, Y. (2024) · 2024
Closest in time.
On pretraining data diversity for self-supervised learning
Hammoud, H. A. A. K., Das, T., Pizzati, F., Torr, P., Bibi, A., and Ghanem, B. (2024) · 2024
Closest in time.
Dinov2: Learning robust visual features without supervision
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al. (2024) · 2024
Closest in time.
Grounded sam: Assembling open-world models for diverse visual tasks
Ren, T., Liu, S., Zeng, A., Lin, J., Li, K., Cao, H., Chen, J., Huang, X., Chen, Y., Yan, F., et al. (2024) · 2024
Closest in time.
Learning vision from models rivals learning vision from data
Tian, Y., Fan, L., Chen, K., Katabi, D., Krishnan, D., and Isola, P. (2024) · 2024
Closest in time.
Real-fake: Effective training data synthesis through distribution matching
Yuan, J., Zhang, J., Sun, S., Torr, P., and Zhao, B. (2024) · 2024
Closest in time.