Fetching the paper…
Reading the bibliography…
The ability to understand visual concepts and replicate and compose these concepts from images is a central goal for computer vision.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Stackgan++: Realistic image synthesis with stacked generative adversarial networks
Zhang, H.; Xu, T.; Li, H.; Zhang, S.; Wang, X.; Huang, X.; and Metaxas, D. N. 2018 · 1962
Earlier work this paper cites.
The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
Kuznetsova, A.; Rom, H.; Alldrin, N.; Uijlings, J.; Krasin, I.; Pont-Tuset, J.; Kamali, S.; Popov, S.; Malloci, M.; Kolesnikov, A.; et al. 2020 · 1981
Earlier work this paper cites.
The big book of concepts
Murphy, G. 2004 · 2004
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
Wah, C.; Branson, S.; Welinder, P.; Perona, P.; and Belongie, S. 2011 · 2011
Earlier work this paper cites.
Deep residual learning for image recognition. arXiv 2015
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
Naeini, M. P.; Cooper, G.; and Hauskrecht, M. 2015 · 2015
Earlier work this paper cites.
On calibration of modern neural networks
Guo, C.; Pleiss, G.; Sun, Y.; and Weinberger, K. Q. 2017 · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; and Hochreiter, S. 2017 · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; et al. 2017 · 2017
Earlier work this paper cites.
Deeper, broader and artier domain generalization
Li, D.; Yang, Y.; Song, Y.-Z.; and Hospedales, T. M. 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A.; Vinyals, O.; et al. 2017 · 2017
Earlier work this paper cites.
Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks
Zhang, H.; Xu, T.; Li, H.; Zhang, S.; Wang, X.; Huang, X.; and Metaxas, D. N. 2017 · 2017
Earlier work this paper cites.
Deep Anomaly Detection with Outlier Exposure
Hendrycks, D.; Mazeika, M.; and Dietterich, T. 2019 · 2019
Earlier work this paper cites.
The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision
Mao, J.; Gan, C.; Kohli, P.; Tenenbaum, J. B.; and Wu, J. 2019 · 2019
Earlier work this paper cites.
Concept bottleneck models
Koh, P. W.; Nguyen, T.; Tang, Y. S.; Mussmann, S.; Pierson, E.; Kim, B.; and Liang, P. 2020 · 2020
Cited alongside, same era.
WeaQA: Weak Supervision via Captions for Visual Question Answering
Banerjee, P.; Gokhale, T.; Yang, Y.; and Baral, C. 2021 · 2021
Cited alongside, same era.
CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Hessel, J.; Holtzman, A.; Forbes, M.; Le Bras, R.; and Choi, Y. 2021 · 2021
Cited alongside, same era.
Vilt: Vision-and-language transformer without convolution or region supervision
Kim, W.; Son, B.; and Kim, I. 2021 · 2021
Cited alongside, same era.
Text-to-image generation grounded by fine-grained user attention
Koh, J. Y.; Baldridge, J.; Lee, H.; and Yang, Y. 2021 · 2021
Cited alongside, same era.
Benchmark for compositional text-to-image synthesis
Park, D. H.; Azadi, S.; Liu, X.; Darrell, T.; and Rohrbach, A. 2021 · 2021
GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
Nichol, A. Q.; Dhariwal, P.; Ramesh, A.; Shyam, P.; Mishkin, P.; Mcgrew, B.; Sutskever, I.; and Chen, M. 2022 · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Later among the works it cites.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Ruiz, N.; Li, Y.; Jampani, V.; Pritch, Y.; Rubinstein, M.; and Aberman, K. 2022 · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E. L.; Ghasemipour, K.; Gontijo Lopes, R.; Karagol Ayan, B.; Salimans, T.; et al. 2022 · 2022
Later among the works it cites.
LAION-5B: An open large-scale dataset for training next generation image-text models
Schuhmann, C.; Beaumont, R.; Vencu, R.; Gordon, C. W.; Wightman, R.; Cherti, M.; Coombes, T.; Katta, A.; Mullis, C.; Wortsman, M.; Schramowski, P.; Kundurthy, S. R.; Crowson, K.; Schmidt, L.; Kaczmarczyk, R.; and Jitsev, J. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Zero-shot text-to-image generation
Ramesh, A.; Pavlov, M.; Goh, G.; Gray, S.; Voss, C.; Radford, A.; Chen, M.; and Sutskever, I. 2021 · 2021
Cited alongside, same era.
Dall-eval: Probing the reasoning skills and social biases of text-to-image generative transformers
Cho, J.; Zala, A.; and Bansal, M. 2022 · 2022
Cited alongside, same era.
Diffedit: Diffusion-based semantic image editing with mask guidance
Couairon, G.; Verbeek, J.; Schwenk, H.; and Cord, M. 2022 · 2022
Cited alongside, same era.
An image is worth one word: Personalizing text-to-image generation using textual inversion
Gal, R.; Alaluf, Y.; Atzmon, Y.; Patashnik, O.; Bermano, A. H.; Chechik, G.; and Cohen-Or, D. 2022 · 2022
Cited alongside, same era.
Benchmarking Spatial Relationships in Text-to-Image Generation
Gokhale, T.; Palangi, H.; Nushi, B.; Vineet, V.; Horvitz, E.; Kamar, E.; Baral, C.; and Yang, Y. 2022 · 2022
Cited alongside, same era.
Later among the works it cites.
Winoground: Probing vision and language models for visio-linguistic compositionality
Thrush, T.; Jiang, R.; Bartolo, M.; Singh, A.; Williams, A.; Kiela, D.; and Ross, C. 2022 · 2022
Later among the works it cites.
Gan inversion: A survey
Xia, W.; Zhang, Y.; Yang, Y.; Xue, J.-H.; Zhou, B.; and Yang, M.-H. 2022 · 2022
Later among the works it cites.
Concept embedding models: Beyond the accuracy-explainability trade-off
Zarlenga, M. E.; Pietro, B.; Gabriele, C.; Giuseppe, M.; Giannini, F.; Diligenti, M.; Zohreh, S.; Frederic, P.; Melacci, S.; Adrian, W.; et al. 2022 · 2022
Later among the works it cites.
Synthetic Data from Diffusion Models Improves ImageNet Classification
Azizi, S.; Kornblith, S.; Saharia, C.; Norouzi, M.; and Fleet, D. J. 2023 · 2023
Closest in time.
PixArt-alpha: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Chen, J.; Yu, J.; Ge, C.; Yao, L.; Xie, E.; Wu, Y.; Wang, Z.; Kwok, J.; Luo, P.; Lu, H.; et al. 2023 · 2023
Closest in time.
Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation
Cho, J.; Li, L.; Yang, Z.; Gan, Z.; Wang, L.; and Bansal, M. 2023 · 2023
Closest in time.
T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation
Huang, K.; Sun, K.; Xie, E.; Li, Z.; and Liu, X. 2023 · 2023
Closest in time.
ECLIPSE: A Resource-Efficient Text-to-Image Prior for Image Generations
Patel, M.; Kim, C.; Cheng, S.; Baral, C.; and Yang, Y. 2023 · 2023
Closest in time.
Multilingual Conceptual Coverage in Text-to-Image Models
Saxon, M.; and Wang, W. Y. 2023 · 2023
Closest in time.
Effective data augmentation with diffusion models
Trabucco, B.; Doherty, K.; Gurinas, M.; and Salakhutdinov, R. 2023 · 2023
Closest in time.