Generative adversarial text to image synthesis
S. Reed, Z. Akata, X. Yan, L. Logeswaran, B. Schiele, and H. Lee · 2016
Earlier work this paper cites.
Inception-v4, inception-resnet and the impact of residual connections on learning
C. Szegedy, S. Ioffe, V. Vanhoucke, and A. Alemi · 2017
Earlier work this paper cites.
Vggface2: A dataset for recognising faces across pose and age
Q. Cao, L. Shen, W. Xie, O. M. Parkhi, and A. Zisserman · 2018
Earlier work this paper cites.
Attngan: Fine-grained text to image generation with attentional generative adversarial networks
T. Xu, P. Zhang, Q. Huang, H. Zhang, Z. Gan, X. Huang, and X. He · 2018
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Ppr10k: A large-scale portrait photo retouching dataset with human-region mask and group-level consistency
J. Liang, H. Zeng, M. Cui, X. Xie, and L. Zhang · 2021
Earlier work this paper cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Original
A. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Earlier work this paper cites.
Encoding in style: a stylegan encoder for image-to-image translation
E. Richardson, Y. Alaluf, O. Patashnik, Y. Nitzan, Y. Azar, S. Shapiro, and D. Cohen-Or · 2021
Earlier work this paper cites.
Hyperextended lightface: A facial attribute analysis framework
S. I. Serengil and A. Ozpinar · 2021
Earlier work this paper cites.
Clip2stylegan: Unsupervised extraction of stylegan edit directions
R. Abdal, P. Zhu, J. Femiani, N. Mitra, and P. Wonka · 2022
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Original
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Earlier work this paper cites.
Blended diffusion for text-driven editing of natural images
O. Avrahami, D. Lischinski, and O. Fried · 2022
Earlier work this paper cites.
ediffi: Text-to-image diffusion models with an ensemble of expert denoisers
Original
Y. Balaji, S. Nah, X. Huang, A. Vahdat, J. Song, K. Kreis, M. Aittala, T. Aila, S. Laine, B. Catanzaro, et al · 2022
Earlier work this paper cites.