Fetching the paper…
Reading the bibliography…
Text-to-image synthesis has recently seen significant progress thanks to large pretrained language models, large-scale training data, and the introduction of scalable model families such as diffusion and autoregressive models.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Generative adversarial networks
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Yfcc100m: The new data in multimedia research
Thomee, B., Shamma, D. A., Friedland, G., Elizalde, B., Ni, K., Poland, D., Borth, D., and Li, L.-J · 2016
Earlier work this paper cites.
GANs trained by a two time-scale update rule converge to a local Nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Hoffer, E., Hubara, I., and Soudry, D · 2017
Earlier work this paper cites.
Geometric Gan
Lim, J. H. and Ye, J. C · 2017
Earlier work this paper cites.
Megapixel size image creation using generative adversarial networks
Marchesi, M · 2017
Earlier work this paper cites.
Progressive growing of GANs for improved quality, stability, and variation
Karras, T., Aila, T., Laine, S., and Lehtinen, J · 2018
Earlier work this paper cites.
cGANs with projection discriminator
Miyato, T. and Koyama, M · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Earlier work this paper cites.
Esrgan: Enhanced super-resolution generative adversarial networks
Wang, X., Yu, K., Wu, S., Gu, J., Liu, Y., Dong, C., Qiao, Y., and Change Loy, C · 2018
Earlier work this paper cites.
Group normalization
Wu, Y. and He, K · 2018
Earlier work this paper cites.
Large scale GAN training for high fidelity natural image synthesis
Brock, A., Donahue, J., and Simonyan, K · 2019
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
Karras, T., Laine, S., and Aila, T · 2019
Earlier work this paper cites.
EfficientNet: Rethinking model scaling for convolutional neural networks
Tan, M. and Le, Q · 2019
Earlier work this paper cites.
Deep polynomial neural networks
Chrysos, G., Moschoglou, S., Bouritsas, G., Deng, J., Panagakis, Y., and Zafeiriou, S · 2020
Earlier work this paper cites.
GANSpace: Discovering interpretable GAN controls
Härkönen, E., Hertzmann, A., Lehtinen, J., and Paris, S · 2020
Earlier work this paper cites.
Analyzing and improving the image quality of StyleGAN
Karras, T., Laine, S., Aittala, M., Hellsten, J., Lehtinen, J., and Aila, T · 2020
Earlier work this paper cites.
Interpreting the latent space of GANs for semantic face editing
Shen, Y., Gu, J., Tang, X., and Zhou, B · 2020
Earlier work this paper cites.
Differentiable augmentation for data-efficient GAN training
Zhao, S., Liu, Z., Lin, J., Zhu, J.-Y., and Han, S · 2020
Earlier work this paper cites.
StyleFlow: Attribute-conditioned exploration of StyleGAN-generated images using conditional continuous normalizing flows
Abdal, R., Zhu, P., Mitra, N. J., and Wonka, P · 2021
Cited alongside, same era.
Deep VIT features as dense visual descriptors
Amir, S., Gandelsman, Y., Bagon, S., and Dekel, T · 2021
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A · 2021
Cited alongside, same era.
Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Changpinyo, S., Sharma, P., Ding, N., and Soricut, R · 2021
Cited alongside, same era.
CoPE: conditional image generation using polynomial expansions
Chrysos, G. G. and Panagakis, Y · 2021
Cited alongside, same era.
Disentangling random and cyclic effects in time-lapse sequences
Härkönen, E., Aittala, M., Kynkäänniemi, T., Laine, S., Aila, T., and Lehtinen, J · 2022
Later among the works it cites.
Classifier-free diffusion guidance
Ho, J. and Salimans, T · 2022
Later among the works it cites.
StyleFusion: A generative model for disentangling spatial segments
Kafri, O., Patashnik, O., Alaluf, Y., and Cohen-Or, D · 2022
Later among the works it cites.
Elucidating the design space of diffusion-based generative models
Karras, T., Aittala, M., Aila, T., and Laine, S · 2022
Later among the works it cites.
The role of ImageNet classes in Fréchet inception distance
Kynkäänniemi, T., Karras, T., Aittala, M., Aila, T., and Lehtinen, J · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
RedCaps: web-curated image-text data created by the people, for the people
Desai, K., Kaul, G., Aysola, Z., and Johnson, J · 2021
Cited alongside, same era.
Diffusion models beat GANs on image synthesis
Dhariwal, P. and Nichol, A · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2021
Cited alongside, same era.
CLIPScore: A reference-free evaluation metric for image captioning
Hessel, J., Holtzman, A., Forbes, M., Bras, R. L., and Choi, Y · 2021
Cited alongside, same era.
Alias-free generative adversarial networks
Karras, T., Aittala, M., Laine, S., Härkönen, E., Hellsten, J., Lehtinen, J., and Aila, T · 2021
Cited alongside, same era.
FuseDream: Training-free text-to-image generation with improved clip+ gan space optimization
Liu, X., Gong, C., Wu, L., Zhang, S., Su, H., and Liu, Q · 2021
Cited alongside, same era.
StyleCLIP: Text-driven manipulation of StyleGAN imagery
Patashnik, O., Wu, Z., Shechtman, E., Cohen-Or, D., and Lischinski, D · 2021
Cited alongside, same era.
All you need is one gpu: Inference benchmark for stable diffusion, 2022
Lambda Labs · 2022
Later among the works it cites.
A convnet for the 2020s
Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S · 2022
Later among the works it cites.
DPM-Solver: A fast ODE solver for diffusion probabilistic model sampling in around 10 steps
Lu, C., Zhou, Y., Bao, F., Chen, J., Li, C., and Zhu, J · 2022
Later among the works it cites.
On distillation of guided diffusion models
Meng, C., Gao, R., Kingma, D. P., Ermon, S., Ho, J., and Salimans, T · 2022
Later among the works it cites.
GLIDE: Towards photorealistic image generation and editing with text-guided diffusion models
Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., and Chen, M · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with CLIP latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Later among the works it cites.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., and Aberman, K · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., et al · 2022
Later among the works it cites.
Progressive distillation for fast sampling of diffusion models
Salimans, T. and Ho, J · 2022
Later among the works it cites.
StyleGAN-XL: Scaling StyleGAN to large diverse datasets
Sauer, A., Schwarz, K., and Geiger, A · 2022
Later among the works it cites.
LAION-5B: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al · 2022
Later among the works it cites.
Flava: A foundational language and vision alignment model
Singh, A., Hu, R., Goswami, V., Couairon, G., Galuba, W., Rohrbach, M., and Kiela, D · 2022
Later among the works it cites.
DF-GAN: A simple and effective baseline for text-to-image synthesis
Tao, M., Tang, H., Wu, F., Jing, X.-Y., Bao, B.-K., and Xu, C · 2022
Later among the works it cites.
Scaling autoregressive models for content-rich text-to-image generation
Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., et al · 2022
Later among the works it cites.
Towards language-free training for text-to-image generation
Zhou, Y., Zhang, R., Chen, C., Li, C., Tensmeyer, C., Yu, T., Gu, J., Xu, J., and Sun, T · 2022
Later among the works it cites.