Fetching the paper…
Reading the bibliography…
When humans read a specific text, they often visualize the corresponding images, and we hope that computers can do the same.
C. J. C. H. Watkins, “Learning from delayed rewards,” 1989
1989
Earlier work this paper cites.
Y. Bengio, R. Ducharme, and P. Vincent, “A neural probabilistic language model,” in
2000
Earlier work this paper cites.
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,”
2004
Earlier work this paper cites.
B. Goertzel and C. Pennachin,
2007
Earlier work this paper cites.
M.-E. Nilsback and A. Zisserman, “Automated flower classification over a large number of classes,” in
2008
Earlier work this paper cites.
D. Marr,
2010
Earlier work this paper cites.
C. Rashtchian, P. Young, M. Hodosh, and J. Hockenmaier, “Collecting image annotations using amazon’s mechanical turk,” in
2010
Earlier work this paper cites.
A. Hore and D. Ziou, “Image quality metrics: Psnr vs. ssim,” in
2010
Earlier work this paper cites.
D. J. Bartholomew, M. Knott, and I. Moustaki,
2011
Earlier work this paper cites.
V. Ordonez, G. Kulkarni, and T. Berg, “Im2text: Describing images using 1 million captioned photographs,” in
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in
2012
Earlier work this paper cites.
D. P. Kingma, “Auto-encoding variational bayes,”
2013
Earlier work this paper cites.
M. Hodosh, P. Young, and J. Hockenmaier, “Framing image description as a ranking task: Data, models and evaluation metrics,”
2013
Earlier work this paper cites.
B. Goertzel, “Artificial general intelligence: concept, state of the art, and future prospects,”
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in
2014
Earlier work this paper cites.
M. Mirza, “Conditional generative adversarial nets,”
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in
2014
Earlier work this paper cites.
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier, “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,”
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,”
2015
Earlier work this paper cites.
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” in
2015
Earlier work this paper cites.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in
2015
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in
2015
Earlier work this paper cites.
B. Romera-Paredes and P. Torr, “An embarrassingly simple approach to zero-shot learning,” in
2015
Earlier work this paper cites.
R. Vedantam, C. Lawrence Zitnick, and D. Parikh, “Cider: Consensus-based image description evaluation,” in
2015
Earlier work this paper cites.
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik, “Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,” in
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “Vqa: Visual question answering,” in
2015
Earlier work this paper cites.
N. Dhanachandra, K. Manglem, and Y. J. Chanu, “Image segmentation using k-means clustering algorithm and subtractive clustering algorithm,”
2015
Earlier work this paper cites.
S. Reed, Z. Akata, X. Yan, L. Logeswaran, B. Schiele, and H. Lee, “Generative adversarial text to image synthesis,” in
2016
Earlier work this paper cites.
J. T. Rolfe, “Discrete variational autoencoders,”
2016
Earlier work this paper cites.
J. Mao, J. Xu, K. Jing, and A. L. Yuille, “Training and evaluating multimodal word embeddings with large-scale web annotated images,” in
2016
Earlier work this paper cites.
X. Yan, J. Yang, K. Sohn, and H. Lee, “Attribute2image: Conditional image generation from visual attributes,” in
2016
Earlier work this paper cites.
A. Van den Oord, N. Kalchbrenner, L. Espeholt, O. Vinyals, A. Graves
2016
Earlier work this paper cites.
D. Lopez-Paz and M. Oquab, “Revisiting classifier two-sample tests,”
2016
Earlier work this paper cites.
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen, “Improved techniques for training gans,” in
2016
Earlier work this paper cites.
H. Zhang, T. Xu, H. Li, S. Zhang, X. Wang, X. Huang, and D. N. Metaxas, “Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks,” in
2017
Earlier work this paper cites.
A. Vaswani, “Attention is all you need,” in
2017
Earlier work this paper cites.
A. Van Den Oord, O. Vinyals
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma
2017
Earlier work this paper cites.
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the v in vqa matter: Elevating the role of image understanding in visual question answering,” in
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
H. Zhang, “mixup: Beyond empirical risk minimization,”
2017
Earlier work this paper cites.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” in
2017
Earlier work this paper cites.
T. Xu, P. Zhang, Q. Huang, H. Zhang, Z. Gan, X. Huang, and X. He, “Attngan: Fine-grained text to image generation with attentional generative adversarial networks,” in
2018
Earlier work this paper cites.
H. Zhang, T. Xu, H. Li, S. Zhang, X. Wang, X. Huang, and D. N. Metaxas, “Stackgan++: Realistic image synthesis with stacked generative adversarial networks,”
2018
Earlier work this paper cites.
A. Radford, “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
N. Bodla, G. Hua, and R. Chellappa, “Semi-supervised fusedgan for conditional image generation,” in
2018
Earlier work this paper cites.
Z. Zhang, Y. Xie, and L. Yang, “Photographic text-to-image synthesis with a hierarchically-nested adversarial network,” in
2018
Earlier work this paper cites.
J. Devlin, “Bert: Pre-training of deep bidirectional transformers for language understanding,”
2018
Earlier work this paper cites.
A. Brock, “Large scale gan training for high fidelity natural image synthesis,”
2018
Earlier work this paper cites.
Z. Liu, P. Luo, X. Wang, and X. Tang, “Large-scale celebfaces attributes (celeba) dataset,”
2018
Earlier work this paper cites.
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in
2018
Earlier work this paper cites.
M. Bińkowski, D. J. Sutherland, M. Arbel, and A. Gretton, “Demystifying mmd gans,”
2018
Earlier work this paper cites.
M. S. Sajjadi, O. Bachem, M. Lucic, O. Bousquet, and S. Gelly, “Assessing generative models via precision and recall,” in
2018
Earlier work this paper cites.
Y. Cui, G. Yang, A. Veit, X. Huang, and S. Belongie, “Learning to evaluate image captioning,” in
2018
Earlier work this paper cites.
S. Kolouri, G. K. Rohde, and H. Hoffmann, “Sliced wasserstein distance for learning gaussian mixture models,” in
2018
Earlier work this paper cites.
M. Zhu, P. Pan, W. Chen, and Y. Yang, “Dm-gan: Dynamic memory generative adversarial networks for text-to-image synthesis,” in
2019
Earlier work this paper cites.
J. Pei, L. Deng, S. Song, M. Zhao, Y. Zhang, S. Wu, G. Wang, Z. Zou, Z. Wu, W. He
2019
Earlier work this paper cites.
T. Qiao, J. Zhang, D. Xu, and D. Tao, “Mirrorgan: Learning text-to-image generation by redescription,” in
2019
Earlier work this paper cites.
H. Tang, D. Xu, N. Sebe, Y. Wang, J. J. Corso, and Y. Yan, “Multi-channel attention selection gan with cascaded semantic guidance for cross-view image translation,” in
2019
Earlier work this paper cites.
M. Cha, Y. L. Gwon, and H. Kung, “Adversarial learning of semantic relevance in text to image synthesis,” in
2019
Earlier work this paper cites.
W. Li, P. Zhang, L. Zhang, Q. Huang, X. He, S. Lyu, and J. Gao, “Object-driven text-to-image synthesis via adversarial training,” in
2019
Earlier work this paper cites.
T. Park, M.-Y. Liu, T.-C. Wang, and J.-Y. Zhu, “Gaugan: semantic image synthesis with spatially adaptive normalization,” in
2019
Earlier work this paper cites.
M. Lee and J. Seok, “Controllable generative adversarial network,”
2019
Earlier work this paper cites.
P. Bachman, R. D. Hjelm, and W. Buchwalter, “Learning representations by maximizing mutual information across views,” in
2019
Earlier work this paper cites.
H. Agrawal, K. Desai, Y. Wang, X. Chen, R. Jain, M. Johnson, D. Batra, D. Parikh, S. Lee, and P. Anderson, “Nocaps: Novel object captioning at scale,” in
2019
Earlier work this paper cites.
R. Zellers, Y. Bisk, A. Farhadi, and Y. Choi, “From recognition to cognition: Visual commonsense reasoning,” in
2019
Earlier work this paper cites.
T. Kynkäänniemi, T. Karras, S. Laine, J. Lehtinen, and T. Aila, “Improved precision and recall metric for assessing generative models,” in
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Cheng, F. Wu, Y. Tian, L. Wang, and D. Tao, “Rifegan: Rich feature generation for text-to-image synthesis from prior knowledge,” in
2020
Earlier work this paper cites.
T. B. Brown, “Language models are few-shot learners,”
2020
Earlier work this paper cites.
A. Dosovitskiy, “An image is worth 16x16 words: Transformers for image recognition at scale,”
2020
Earlier work this paper cites.
J. Liang, W. Pei, and F. Lu, “Cpgan: Content-parsing generative adversarial networks for text-to-image synthesis,” in
2020
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in
2020
Earlier work this paper cites.
H. Tang, D. Xu, Y. Yan, P. H. Torr, and N. Sebe, “Local class-specific and global image-level generative adversarial networks for semantic-guided scene generation,” in
2020
Earlier work this paper cites.
D. Croce, G. Castellucci, and R. Basili, “Gan-bert: Generative adversarial learning for robust text classification with a bunch of labeled examples,” in
2020
Earlier work this paper cites.
M. Chen, A. Radford, R. Child, J. Wu, H. Jun, D. Luan, and I. Sutskever, “Generative pretraining from pixels,” in
2020
Earlier work this paper cites.
B. Zhu and C.-W. Ngo, “Cookgan: Causality based text-to-image synthesis,” in
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
V. Sitzmann, J. Martel, A. Bergman, D. Lindell, and G. Wetzstein, “Implicit neural representations with periodic activation functions,” in
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
T. Karras, S. Laine, M. Aittala, J. Hellsten, J. Lehtinen, and T. Aila, “Analyzing and improving the image quality of stylegan,” in
2020
Earlier work this paper cites.
H. Lee, S. Yoon, F. Dernoncourt, D. S. Kim, T. Bui, and K. Jung, “Vilbertscore: Evaluating image caption using vision-and-language bert,” in
2020
Earlier work this paper cites.
T. Hinz, S. Heinrich, and S. Wermter, “Semantic object accuracy for generative text-to-image synthesis,”
2020
Earlier work this paper cites.
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” in
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
M. Ding, Z. Yang, W. Hong, W. Zheng, C. Zhou, D. Yin, J. Lin, X. Zou, Z. Shao, H. Yang
2021
Earlier work this paper cites.
P. Esser, R. Rombach, and B. Ommer, “Taming transformers for high-resolution image synthesis,” in
2021
Earlier work this paper cites.
S. Ruan, Y. Zhang, K. Zhang, Y. Fan, F. Tang, Q. Liu, and E. Chen, “Dae-gan: Dynamic aspect-aware gan for text-to-image synthesis,” in
2021
Earlier work this paper cites.
Z. Niu, G. Zhong, and H. Yu, “A review on the attention mechanism of deep learning,”
2021
Earlier work this paper cites.
S. Frolov, T. Hinz, F. Raue, J. Hees, and A. Dengel, “Adversarial text-to-image synthesis: A review,”
2021
Earlier work this paper cites.
R. Zhou, C. Jiang, and Q. Xu, “A survey on generative adversarial network-based text-to-image synthesis,”
2021
Earlier work this paper cites.
H. Zhang, J. Y. Koh, J. Baldridge, H. Lee, and Y. Yang, “Cross-modal contrastive learning for text-to-image generation,” in
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark
2021
Earlier work this paper cites.
P. Dhariwal and A. Nichol, “Diffusion models beat gans on image synthesis,” in
2021
Earlier work this paper cites.
H. Wang, G. Lin, S. C. Hoi, and C. Miao, “Cycle-consistent inverse gan for text-to-image synthesis,” in
2021
Earlier work this paper cites.
Y. Qiao, Q. Chen, C. Deng, N. Ding, Y. Qi, M. Tan, X. Ren, and Q. Wu, “R-gan: Exploring human-like way for reasonable text-to-image synthesis via generative adversarial networks,” in
2021
Earlier work this paper cites.
J. Lin, R. Men, A. Yang, C. Zhou, M. Ding, Y. Zhang, P. Wang, A. Wang, L. Jiang, X. Jia
2021
Earlier work this paper cites.
X. Wu, Z. Hu, L. Sheng, and D. Xu, “Styleformer: Real-time arbitrary style transfer via parametric style composition,” in
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Z. Zhang, J. Ma, C. Zhou, R. Men, Z. Li, M. Ding, J. Tang, J. Zhou, and H. Yang, “Ufc-bert: Unifying multi-modal controls for conditional image synthesis,”
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
S. Changpinyo, P. Sharma, N. Ding, and R. Soricut, “Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts,” in
2021
Earlier work this paper cites.
C.-B. Zhang, P.-T. Jiang, Q. Hou, Y. Wei, Q. Han, Z. Li, and M.-M. Cheng, “Delving deep into label smoothing,”
2021
Earlier work this paper cites.
C.-Y. Bai, H.-T. Lin, C. Raffel, and W. C.-w. Kan, “On training sample memorization: Lessons from benchmarking generative modeling with a large-scale competition,” in
2021
Earlier work this paper cites.
D. H. Park, S. Azadi, X. Liu, T. Darrell, and A. Rohrbach, “Benchmark for compositional text-to-image synthesis,” in
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
M. A. H. Palash, M. A. Al Nasim, A. Dhali, and F. Afrin, “Fine-grained image generation from bangla text description using attentional generative adversarial network,” in
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
M. Ding, W. Zheng, W. Hong, and J. Tang, “Cogview2: Faster and better text-to-image generation via hierarchical transformers,” in
2022
Earlier work this paper cites.
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans
2022
Earlier work this paper cites.
K. Crowson, S. Biderman, D. Kornis, D. Stander, E. Hallahan, L. Castricato, and E. Raff, “Vqgan-clip: Open domain image generation and editing with natural language guidance,” in
2022
Earlier work this paper cites.
S. Gu, D. Chen, J. Bao, F. Wen, B. Zhang, D. Chen, L. Yuan, and B. Guo, “Vector quantized diffusion model for text-to-image synthesis,” in
2022
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in
2022
Earlier work this paper cites.
M. Tao, H. Tang, F. Wu, X.-Y. Jing, B.-K. Bao, and C. Xu, “Df-gan: A simple and effective baseline for text-to-image synthesis,” in
2022
Earlier work this paper cites.
O. Gafni, A. Polyak, O. Ashual, S. Sheynin, D. Parikh, and Y. Taigman, “Make-a-scene: Scene-based text-to-image generation with human priors,” in
2022
Earlier work this paper cites.
C. Wu, J. Liang, L. Ji, F. Yang, Y. Fang, D. Jiang, and N. Duan, “Nüwa: Visual synthesis pre-training for neural visual world creation,” in
2022
Earlier work this paper cites.
T. M. Dinh, R. Nguyen, and B.-S. Hua, “Tise: Bag of metrics for text-to-image synthesis evaluation,” in
2022
Earlier work this paper cites.
J. Ho and T. Salimans, “Classifier-free diffusion guidance,”
2022
Earlier work this paper cites.
Q. Cheng, K. Wen, and X. Gu, “Vision-language matching for text-to-image synthesis via generative adversarial networks,”
2022
Earlier work this paper cites.
W. Liao, K. Hu, M. Y. Yang, and B. Rosenhahn, “Text to image generation with semantic-spatial aware gan,” 2022
2022
Earlier work this paper cites.
Y. Zhou, R. Zhang, C. Chen, C. Li, C. Tensmeyer, T. Yu, J. Gu, J. Xu, and T. Sun, “Towards language-free training for text-to-image generation,” 2022
2022
Earlier work this paper cites.
F. Wu, L. Liu, F. Hao, F. He, and J. Cheng, “Text-to-image synthesis based on object-guided joint-decoding transformer,” in
2022
Earlier work this paper cites.
O. Avrahami, D. Lischinski, and O. Fried, “Blended diffusion for text-driven editing of natural images,” in
2022
Earlier work this paper cites.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Z. Xu, T. Lin, H. Tang, F. Li, D. He, N. Sebe, R. Timofte, L. Van Gool, and E. Ding, “Predict, prevent, and evaluate: Disentangled text-driven image manipulation empowered by pre-trained vision-language model,” in
2022
Earlier work this paper cites.
J. Lezama, H. Chang, L. Jiang, and I. Essa, “Improved masked image generation with token-critic,” in
2022
Earlier work this paper cites.
T. Salimans and J. Ho, “Progressive distillation for fast sampling of diffusion models,”
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
G. Kim, T. Kwon, and J. C. Ye, “Diffusionclip: Text-guided diffusion models for robust image manipulation,” in
2022
Earlier work this paper cites.
X. Wu, H. Zhao, L. Zheng, S. Ding, and X. Li, “Adma-gan: Attribute-driven memory augmented gans for text-to-image generation.” in
2022
Earlier work this paper cites.
Z. Shi, Z. Chen, Z. Xu, W. Yang, and L. Huang, “Athom: Two divergent attentions stimulated by homomorphic training in text-to-image synthesis,” in
2022
Earlier work this paper cites.
Z. Chen, Z. Mao, S. Fang, and B. Hu, “Background layout generation and object knowledge transfer for text-to-image generation,” in
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Z. Li, M. R. Min, K. Li, and C. Xu, “Stylet2i: Toward compositional and high-fidelity text-to-image synthesis,” 2022
2022
Earlier work this paper cites.
K. Yan, L. Ji, C. Wu, J. Bao, M. Zhou, N. Duan, and S. Ma, “Trace controlled text to image generation,” in
2022
Earlier work this paper cites.
Y. Jiang, S. Yang, H. Qiu, W. Wu, C. C. Loy, and Z. Liu, “Text2human: Text-driven controllable human image generation,”
2022
Earlier work this paper cites.
A. Maharana, D. Hannan, and M. Bansal, “Storydall-e: Adapting pretrained text-to-image transformers for story continuation,” in
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
M. Huang, Z. Mao, P. Wang, Q. Wang, and Y. Zhang, “Dse-gan: Dynamic semantic evolution generative adversarial network for text-to-image generation,” in
2022
Earlier work this paper cites.
L. H. Li, P. Zhang, H. Zhang, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J.-N. Hwang
2022
Earlier work this paper cites.
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman
2022
Cited alongside, same era.
2022
Cited alongside, same era.
T. Thrush, R. Jiang, M. Bartolo, A. Singh, A. Williams, D. Kiela, and C. Ross, “Winoground: Probing vision and language models for visio-linguistic compositionality,” in
2022
Cited alongside, same era.
2022
Cited alongside, same era.
J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in
G. Xiao, T. Yin, W. T. Freeman, F. Durand, and S. Han, “Fastcomposer: Tuning-free multi-subject image generation with localized attention,”
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
P. Li, Q. Nie, Y. Chen, X. Jiang, K. Wu, Y. Lin, Y. Liu, J. Peng, C. Wang, and F. Zheng, “Tuning-free image customization with image and text guidance,” in
2024
Closest in time.
Y. Cai, Y. Wei, Z. Ji, J. Bai, H. Han, and W. Zuo, “Decoupled textual embeddings for customized image generation,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
2022
Cited alongside, same era.
J.-H. Kim, Y. Kim, J. Lee, K. M. Yoo, and S.-W. Lee, “Mutual information divergence: A unified metric for multimodal generative models,” 2022
2022
Cited alongside, same era.
L. Zhang, Y. Zhou, C. Barnes, S. Amirghodsi, Z. Lin, E. Shechtman, and J. Shi, “Perceptual artifacts localization for inpainting,” in
2022
Cited alongside, same era.
L. Zhang, A. Rao, and M. Agrawala, “Adding conditional control to text-to-image diffusion models,” in
2023
Cited alongside, same era.
2023
Cited alongside, same era.
J. Wu, W. Gan, Z. Chen, S. Wan, and H. Lin, “Ai-generated content (aigc): A survey,”
2023
Cited alongside, same era.
M. Tao, B.-K. Bao, H. Tang, and C. Xu, “Galip: Generative adversarial clips for text-to-image synthesis,” in
2023
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
S. Cui, J. Guo, X. An, J. Deng, Y. Zhao, X. Wei, and Z. Feng, “Idadapter: Learning mixed features for tuning-free personalization of text-to-image models,” in
2024
Closest in time.
Y. Zhang, Y. Song, J. Liu, R. Wang, J. Yu, H. Tang, H. Li, X. Tang, Y. Hu, H. Pan
2024
Closest in time.
C. Mou, X. Wang, L. Xie, Y. Wu, J. Zhang, Z. Qi, and Y. Shan, “T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models,” in
2024
Closest in time.
J. Kim, J. Park, and W. Rhee, “Selectively informative description can reduce undesired embedding entanglements in text-to-image personalization,” in
2024
Closest in time.
Y. Zeng, V. M. Patel, H. Wang, X. Huang, T.-C. Wang, M.-Y. Liu, and Y. Balaji, “Jedi: Joint-image diffusion models for finetuning-free personalized text-to-image generation,” in
2024
Closest in time.
J. Nam, H. Kim, D. Lee, S. Jin, S. Kim, and S. Chang, “Dreammatcher: Appearance matching self-attention for semantically-consistent text-to-image personalization,” in
2024
Closest in time.
X. Li, X. Hou, and C. C. Loy, “When stylegan meets stable diffusion: a w+ adapter for personalized image generation,” in
2024
Closest in time.
S. Zhao, D. Chen, Y.-C. Chen, J. Bao, S. Hao, L. Yuan, and K.-Y. K. Wong, “Uni-controlnet: All-in-one control to text-to-image diffusion models,” in
2024
Closest in time.
D. Li, J. Li, and S. Hoi, “Blip-diffusion: Pre-trained subject representation for controllable text-to-image generation and editing,” in
2024
Closest in time.
K. Mei, M. Delbracio, H. Talebi, Z. Tu, V. M. Patel, and P. Milanfar, “Codi: Conditional diffusion distillation for higher-fidelity and faster image generation,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Y. Wang, R. Gao, K. Chen, K. Zhou, Y. Cai, L. Hong, Z. Li, L. Jiang, D.-Y. Yeung, Q. Xu
2024
Closest in time.
H. Cai, M. Li, Q. Zhang, M.-Y. Liu, and S. Han, “Condition-aware neural network for controlled image generation,” in
2024
Closest in time.
J. Ren, M. Xu, J.-C. Wu, Z. Liu, T. Xiang, and A. Toisoul, “Move anything with layered scene diffusion,” in
2024
Closest in time.
M. Ohanyan, H. Manukyan, Z. Wang, S. Navasardyan, and H. Shi, “Zero-painter: Training-free layout control for text-to-image synthesis,” in
2024
Closest in time.
S. Mo, F. Mu, K. H. Lin, Y. Liu, B. Guan, Y. Li, and B. Zhou, “Freecontrol: Training-free spatial control of any text-to-image diffusion model with any condition,” in
2024
Closest in time.
M. Li, T. Yang, H. Kuang, J. Wu, Z. Wang, X. Xiao, and C. Chen, “Controlnet++ improving conditional controls with efficient consistency feedback,” in
2024
Closest in time.
2024
Closest in time.
D.-Y. Chen, H. Tennent, and C.-W. Hsu, “Artadapter: Text-to-image style transfer using multi-level style encoder and explicit adaptation,” in
2024
Closest in time.
J. Shi, W. Xiong, Z. Lin, and H. J. Jung, “Instantbooth: Personalized text-to-image generation without test-time finetuning,” in
2024
Closest in time.
H. Cho, J. Lee, S. Chang, and Y. Jeong, “One-shot structure-aware stylized image synthesis,” in
2024
Closest in time.
T. Qi, S. Fang, Y. Wu, H. Xie, J. Liu, L. Chen, Q. He, and Y. Zhang, “Deadiff: An efficient stylization diffusion model with disentangled representations,” in
2024
Closest in time.
2024
Closest in time.
A. Hertz, A. Voynov, S. Fruchter, and D. Cohen-Or, “Style aligned image generation via shared attention,” in
2024
Closest in time.
G. Ding, C. Zhao, W. Wang, Z. Yang, Z. Liu, H. Chen, and C. Shen, “Freecustom: Tuning-free customized image generation for multi-concept composition,” in
2024
Closest in time.
Y. Zhang, M. Yang, Q. Zhou, and Z. Wang, “Attention calibration for disentangled text-to-image personalization,” in
2024
Closest in time.
K. Sueyoshi and T. Matsubara, “Predicated diffusion: Predicate logic-based attention guidance for text-to-image diffusion models,” in
2024
Closest in time.
M. Brack, F. Friedrich, K. Kornmeier, L. Tsaban, P. Schramowski, K. Kersting, and A. Passos, “Ledits++: Limitless image editing using text-to-image models,” in
2024
Closest in time.
S. Mahajan, T. Rahman, K. M. Yi, and L. Sigal, “Prompting hard or hardly prompting: Prompt inversion for text-to-image diffusion models,” in
2024
Closest in time.
2024
Closest in time.
Z. Ma, G. Jia, and B. Zhou, “Adapedit: Spatio-temporal guided adaptive editing algorithm for text-based continuity-sensitive image editing,” in
2024
Closest in time.
Z. Yu, H. Li, F. Fu, X. Miao, and B. Cui, “Accelerating text-to-image editing via cache-enabled sparse diffusion inference,” in
2024
Closest in time.
Y. Qiao, F. Wang, J. Su, Y. Zhang, Y. Yu, S. Wu, and G.-J. Qi, “Baret: Balanced attention based real image editing driven by target-text inversion,” in
2024
Closest in time.
2024
Closest in time.
Z. Wang, Y. Huang, D. Song, L. Ma, and T. Zhang, “Promptcharm: Text-to-image generation through multi-modal prompting and refinement,” in
2024
Closest in time.
Z. Xue, G. Song, Q. Guo, B. Liu, Z. Zong, Y. Liu, and P. Luo, “Raphael: Text-to-image generation via large mixture of diffusion paths,” in
2024
Closest in time.
B. Liu, C. Wang, T. Cao, K. Jia, and J. Huang, “Towards understanding cross and self-attention in stable diffusion for text-guided image editing,” in
2024
Closest in time.
Q. Guo and T. Lin, “Focus on your instruction: Fine-grained and multi-instruction image editing by attention modulation,” in
2024
Closest in time.
C. Mou, X. Wang, J. Song, Y. Shan, and J. Zhang, “Diffeditor: Boosting accuracy and flexibility on diffusion-based image editing,” in
2024
Closest in time.
R. Bodur, B. Bhattarai, and T.-K. Kim, “Prompt augmentation for self-supervised text-guided image manipulation,” in
2024
Closest in time.
2024
Closest in time.
M. Patel, C. Kim, S. Cheng, C. Baral, and Y. Yang, “Eclipse: A resource-efficient text-to-image prior for image generations,” in
2024
Closest in time.
S. Liu, W. Yu, Z. Tan, and X. Wang, “Linfusion: 1 gpu, 1 minute, 16k image,”
2024
Closest in time.
2024
Closest in time.
T. H. Nguyen and A. Tran, “Swiftbrush: One-step text-to-image diffusion model with variational score distillation,” in
2024
Closest in time.
S. X. Chen, Y. Vaxman, E. Ben Baruch, D. Asulin, A. Moreshet, K.-C. Lien, M. Sra, and P. Sen, “Tino-edit: Timestep and noise optimization for robust diffusion-based image editing,” in
2024
Closest in time.
H. Li, Y. Zou, Y. Wang, O. Majumder, Y. Xie, R. Manmatha, A. Swaminathan, Z. Tu, S. Ermon, and S. Soatto, “On the scalability of diffusion-based text-to-image generation,” in
2024
Closest in time.
S. Jayasumana, D. Glasner, S. Ramalingam, A. Veit, A. Chakrabarti, and S. Kumar, “Markovgen: Structured prediction for efficient text-to-image generation,” in
2024
Closest in time.
W. Mo, T. Zhang, Y. Bai, B. Su, J.-R. Wen, and Q. Yang, “Dynamic prompt optimizing for text-to-image generation,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Y. Hao, Z. Chi, L. Dong, and F. Wei, “Optimizing prompts for text-to-image generation,” in
2024
Closest in time.
2024
Closest in time.
H. Li, C. Shen, P. Torr, V. Tresp, and J. Gu, “Self-discovering interpretable diffusion latent directions for responsible text-to-image generation,” in
2024
Closest in time.
2024
Closest in time.
H. Zhang, Y. He, and H. Chen, “Steerdiff: Steering towards safe text-to-image diffusion models,”
2024
Closest in time.
Y. Liu, C. Fan, Y. Dai, X. Chen, P. Zhou, and L. Sun, “Metacloak: Preventing unauthorized subject-driven text-to-image diffusion-based synthesis via meta-learning,” in
2024
Closest in time.
M. D’Incà, E. Peruzzo, M. Mancini, D. Xu, V. Goel, X. Xu, Z. Wang, H. Shi, and N. Sebe, “Openbias: Open-set bias detection in text-to-image generative models,” in
2024
Closest in time.
M. Lyu, Y. Yang, H. Hong, H. Chen, X. Jin, Y. He, H. Xue, J. Han, and G. Ding, “One-dimensional adapter to rule them all: Concepts diffusion models and erasing applications,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
H. Liu, Z. Sun, and Y. Mu, “Countering personalized text-to-image generation with influence watermarks,” in
2024
Closest in time.
X. Li, Q. Shen, and K. Kawaguchi, “Va3: Virtually assured amplification attack on probabilistic copyright protection for text-to-image generative models,” in
2024
Closest in time.
F. Wang, Z. Tan, T. Wei, Y. Wu, and Q. Huang, “Simac: A simple anti-customization method for protecting face privacy against text-to-image synthesis of diffusion models,” in
2024
Closest in time.
Q. Phung, S. Ge, and J.-B. Huang, “Grounded text-to-image synthesis with attention refocusing,” in
2024
Closest in time.
L. Qu, W. Wang, Y. Li, H. Zhang, L. Nie, and T.-S. Chua, “Discriminative probing and tuning for text-to-image generation,” in
2024
Closest in time.
D. Zhou, Y. Li, F. Ma, X. Zhang, and Y. Yang, “Migc: Multi-instance generation controller for text-to-image synthesis,” in
2024
Closest in time.
X. Guo, J. Liu, M. Cui, J. Li, H. Yang, and D. Huang, “Initno: Boosting text-to-image diffusion models via initial noise optimization,” in
2024
Closest in time.
C. Liu, X. Li, and H. Ding, “Referring image editing: Object-level image editing via referring expressions,” in
2024
Closest in time.
Q. Zhangli, J. Jiang, D. Liu, L. Yu, X. Dai, A. Ramchandani, G. Pang, D. N. Metaxas, and P. Krishnan, “Layout-agnostic scene text image synthesis with diffusion models,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
R. Rassin, E. Hirsch, D. Glickman, S. Ravfogel, Y. Goldberg, and G. Chechik, “Linguistic binding in diffusion models: Enhancing attribute correspondence through attention map alignment,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
X. Song, J. Cui, H. Zhang, J. Chen, R. Hong, and Y.-G. Jiang, “Doubly abductive counterfactual inference for text-based image editing,” in
2024
Closest in time.
R. Wang, Z. Chen, C. Chen, J. Ma, H. Lu, and X. Lin, “Compositional text-to-image synthesis with attention map control of diffusion models,” in
2024
Closest in time.
H. Nam, G. Kwon, G. Y. Park, and J. C. Ye, “Contrastive denoising score for text-guided latent diffusion image editing,” in
2024
Closest in time.
2024
Closest in time.
P. Ling, L. Chen, P. Zhang, H. Chen, Y. Jin, and J. Zheng, “Freedrag: Feature dragging for reliable point-based image editing,” in
2024
Closest in time.
J. Lu, X. Li, and K. Han, “Regiondrag: Fast region-based image editing with diffusion models,”
2024
Closest in time.
2024
Closest in time.
P.-D. Tudosiu, Y. Yang, S. Zhang, F. Chen, S. McDonagh, G. Lampouras, I. Iacobacci, and S. Parisot, “Mulan: A multi layer annotated dataset for controllable text-to-image generation,” in
2024
Closest in time.
Z. Lv, Y. Wei, W. Zuo, and K.-Y. K. Wong, “Place: Adaptive layout-semantic fusion for semantic image synthesis,” in
2024
Closest in time.
M. Chen, I. Laina, and A. Vedaldi, “Training-free layout control with cross-attention guidance,” in
2024
Closest in time.
C. Jia, M. Luo, Z. Dang, G. Dai, X. Chang, M. Wang, and J. Wang, “Ssmg: Spatial-semantic map guided diffusion model for free-form layout-to-image generation,” in
2024
Closest in time.
S. Li, J. Fu, K. Liu, W. Wang, K.-Y. Lin, and W. Wu, “Cosmicman: A text-to-image foundation model for humans,” in
2024
Closest in time.
Y. Zhou, R. Zhang, J. Gu, and T. Sun, “Customization assistant for text-to-image generation,” in
2024
Closest in time.
S. Narasimhaswamy, U. Bhattacharya, X. Chen, I. Dasgupta, S. Mitra, and M. Hoai, “Handiffuser: Text-to-image generation with realistic hand appearances,” in
2024
Closest in time.
J. Yang, J. Feng, and H. Huang, “Emogen: Emotional image content generation with text-to-image diffusion models,” in
2024
Closest in time.
J. Wang, Z. Sun, Z. Tan, X. Chen, W. Chen, H. Li, C. Zhang, and Y. Song, “Towards effective usage of human-centric priors in diffusion models for text-based human image generation,” in
2024
Closest in time.
C. Zhang, Q. Wu, C. C. Gambardella, X. Huang, D. Phung, W. Ouyang, and J. Cai, “Taming stable diffusion for text to 360 panorama image generation,” in
2024
Closest in time.
C. Liu, H. Wu, Y. Zhong, X. Zhang, Y. Wang, and W. Xie, “Intelligent grimm-open-ended visual storytelling via latent diffusion models,” in
2024
Closest in time.
2024
Closest in time.
X. Xing, H. Zhou, C. Wang, J. Zhang, D. Xu, and Q. Yu, “Svgdreamer: Text guided svg generation with diffusion model,” in
2024
Closest in time.
2024
Closest in time.
T. Wang and M. Ye, “Texfit: Text-driven fashion image editing with diffusion models,” in
2024
Closest in time.
2024
Closest in time.
T.-Y. Cheng, M. Gadelha, T. Groueix, M. Fisher, R. Mech, A. Markham, and N. Trigoni, “Learning continuous 3d words for text-to-image generation,” in
2024
Closest in time.
J. T. Hoe, X. Jiang, C. S. Chan, Y.-P. Tan, and W. Hu, “Interactdiffusion: Interaction control in text-to-image diffusion models,” in
2024
Closest in time.
R. Parihar, V. Sachidanand, S. Mani, T. Karmali, and R. Venkatesh Babu, “Precisecontrol: Enhancing text-to-image diffusion models with fine-grained attribute control,” in
2024
Closest in time.
2024
Closest in time.
P. Zhang, H. Yin, C. Li, and X. Xie, “Tackling the singularities at the endpoints of time intervals in diffusion models,” in
2024
Closest in time.
Y. Lu, M. Zhang, A. J. Ma, X. Xie, and J. Lai, “Coarse-to-fine latent diffusion for pose-guided person image synthesis,” in
2024
Closest in time.
T. Shirakawa and S. Uchida, “Noisecollage: A layout-aware text-to-image diffusion model based on noise cropping and merging,” in
2024
Closest in time.
G. Kwon, S. Jenni, D. Li, J.-Y. Lee, J. C. Ye, and F. C. Heilbron, “Concept weaver: Enabling multi-concept fusion in text-to-image models,” in
2024
Closest in time.
2024
Closest in time.
L. Zhang and M. Agrawala, “Transparent image layer diffusion using latent transparency,”
2024
Closest in time.
T. H. S. Meral, E. Simsar, F. Tombari, and P. Yanardag, “Conform: Contrast is all you need for high-fidelity text-to-image diffusion models,” in
2024
Closest in time.
W. Feng, W. Zhu, T.-j. Fu, V. Jampani, A. Akula, X. He, S. Basu, X. E. Wang, and W. Y. Wang, “Layoutgpt: Compositional visual planning and generation with large language models,” in
2024
Closest in time.
2024
Closest in time.
Y. Huang, L. Xie, X. Wang, Z. Yuan, X. Cun, Y. Ge, J. Zhou, C. Dong, R. Huang, R. Zhang
2024
Closest in time.
Y. Yao, C.-F. Hsu, J.-H. Lin, H. Xie, T. Lin, Y.-N. Huang, H.-H. Shuai, and W.-H. Cheng, “The fabrication of reality and fantasy: Scene generation with llm-assisted prompt interpretation,” in
2024
Closest in time.
J. Liao, X. Chen, Q. Fu, L. Du, X. He, X. Wang, S. Han, and D. Zhang, “Text-to-image generation for abstract concepts,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
P. Grimal, H. Le Borgne, O. Ferret, and J. Tourille, “Tiam-a metric for evaluating alignment in text-to-image generation,” in
2024
Closest in time.
Y. Wang, W. Zhang, J. Zheng, and C. Jin, “High-fidelity person-centric subject-to-image synthesis,” in
2024
Closest in time.
Y. Lin, Y.-W. Chen, Y.-H. Tsai, L. Jiang, and M.-H. Yang, “Text-driven image editing via learnable regions,” in
2024
Closest in time.
2024
Closest in time.
Z. Miao, J. Wang, Z. Wang, Z. Yang, L. Wang, Q. Qiu, and Z. Liu, “Training diffusion models towards diverse image generation with reinforcement learning,” in
2024
Closest in time.
J. Cho, A. Zala, and M. Bansal, “Visual programming for step-by-step text-to-image generation and evaluation,” in
2024
Closest in time.
S. Hartwig, D. Engel, L. Sick, H. Kniesel, T. Payer, T. Ropinski
2024
Closest in time.
S. Ye and F. Liu, “Data extrapolation for text-to-image generation on small datasets,”
2024
Closest in time.
Z. Tan, X. Yang, and K. Huang, “Semantic-aware data augmentation for text-to-image synthesis,” in
2024
Closest in time.
T. Lee, M. Yasunaga, C. Meng, Y. Mai, J. S. Park, A. Gupta, Y. Zhang, D. Narayanan, H. Teufel, M. Bellagente
2024
Closest in time.
2024
Closest in time.
M. Yarom, Y. Bitton, S. Changpinyo, R. Aharoni, J. Herzig, O. Lang, E. Ofek, and I. Szpektor, “What you see is what you read? improving text-image alignment evaluation,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
B. Gordon, Y. Bitton, Y. Shafir, R. Garg, X. Chen, D. Lischinski, D. Cohen-Or, and I. Szpektor, “Mismatch quest: Visual and textual feedback for image-text misalignment,” in
2024
Closest in time.
K. Team, “Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis,” 2024
2024
Closest in time.
2024
Closest in time.
A. Jha, V. Prabhakaran, R. Denton, S. Laszlo, S. Dave, R. Qadri, C. Reddy, and S. Dev, “Visage: A global-scale analysis of visual stereotypes in text-to-image generation,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Y. Sun, X. Li, K. Dalal, J. Xu, A. Vikram, G. Zhang, Y. Dubois, X. Chen, X. Wang, S. Koyejo
2024
Closest in time.
A. Chatterjee, G. B. M. Stan, E. Aflalo, S. Paul, D. Ghosh, T. Gokhale, L. Schmidt, H. Hajishirzi, V. Lal, C. Baral
2025
Closest in time.
S. Kim, S. Jung, B. Kim, M. Choi, J. Shin, and J. Lee, “Safeguard text-to-image diffusion models with human feedback inversion,” in
2025
Closest in time.