Fetching the paper…
Reading the bibliography…
Diffusion-driven text-to-image (T2I) generation has achieved remarkable advancements in recent years.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Earlier work this paper cites.
Tell, draw, and repeat: Generating and modifying images based on continual linguistic instruction
A. El-Nouby, S. Sharma, H. Schulz, D. Hjelm, L. El Asri, S. Ebrahimi Kahou, Y. Bengio, and G. W. Taylor · 2019
Earlier work this paper cites.
Sscr: Iterative language-based image editing via self-supervised counterfactual reasoning
T.-J. Fu, X. E. Wang, S. Grafton, M. Eckstein, and W. Y. Wang · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
Zero-Shot Text-to-Image Generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Earlier work this paper cites.
ediff-i: Text-to-image diffusion models with an ensemble of expert denoisers
Y. Balaji, S. Nah, X. Huang, A. Vahdat, J. Song, Q. Zhang, K. Kreis, M. Aittala, T. Aila, S. Laine, B. Catanzaro, T. Karras, and M.-Y. Liu · 2022
Earlier work this paper cites.
Cogview2: Faster and better text-to-image generation via hierarchical transformers
M. Ding, W. Zheng, W. Hong, and J. Tang · 2022
Earlier work this paper cites.
Make-a-Scene: Scene-Based Text-to-Image Generation with Human Priors
O. Gafni, A. Polyak, O. Ashual, S. Sheynin, D. Parikh, and Y. Taigman · 2022
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2022
Earlier work this paper cites.
GLIDE: Towards photorealistic image generation and editing with text-guided diffusion models
A. Q. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen · 2022
Earlier work this paper cites.
Hierarchical text-conditional image generation with clip latents
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Earlier work this paper cites.
High-resolution Image Synthesis with Latent Diffusion Models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Earlier work this paper cites.
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. Denton, S. K. Seyed Ghasemipour, B. K. Ayan, S. S. Mahdavi, R. G. Lopes, T. Salimans, J. Ho, D. J. Fleet, and M. Norouzi · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
Scaling Autoregressive Models for Content-Rich Text-to-Image Generation
J. Yu, Y. Xu, J. Y. Koh, T. Luong, G. Baid, Z. Wang, V. Vasudevan, A. Ku, Y. Yang, B. K. Ayan, B. Hutchinson, H. Wei, Z. Parekh, X. Li, H. Zhang, J. Baldridge, and Y. Wu · 2022
Cited alongside, same era.
Spatext: Spatio-textual representation for controllable image generation
O. Avrahami, T. Hayes, O. Gafni, S. Gupta, Y. Taigman, D. Parikh, D. Lischinski, O. Fried, and X. Yin · 2023
Cited alongside, same era.
HRS-Bench: Holistic, Reliable and Scalable Benchmark for Text-to-Image Models
E. M. Bakr, P. Sun, X. Shen, F. F. Khan, L. E. Li, and M. Elhoseiny · 2023
Cited alongside, same era.
Multidiffusion: Fusing diffusion paths for controlled image generation
O. Bar-Tal, L. Yariv, Y. Lipman, and T. Dekel · 2023
Cited alongside, same era.
Mixture of diffusers for scene composition and high-resolution image generation
A. Barbero Jiménez · 2023
Cited alongside, same era.
Layoutllm-t2i: Eliciting layout guidance from llm for text-to-image generation
L. Qu, S. Wu, H. Fei, L. Nie, and T.-S. Chua · 2023
Later among the works it cites.
Stylegan-t: unlocking the power of gans for fast large-scale text-to-image synthesis
A. Sauer, T. Karras, S. Laine, A. Geiger, and T. Aila · 2023
Later among the works it cites.
Harnessing the spatial-temporal attention of diffusion models for high-fidelity text-to-image synthesis
Q. Wu, Y. Liu, H. Zhao, T. Bui, Z. Lin, Y. Zhang, and S. Chang · 2023
Later among the works it cites.
Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion
J. Xie, Y. Li, Y. Huang, H. Liu, W. Zhang, Y. Zheng, and M. Z. Shou · 2023
Later among the works it cites.
Reco: Region-controlled text-to-image generation
Z. Yang, J. Wang, Z. Gan, L. Li, K. Lin, C. Wu, N. Duan, Z. Liu, C. Liu, M. Zeng, and L. Wang · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Muse: Text-to-image generation via masked generative transformers
H. Chang, H. Zhang, J. Barber, A. Maschinot, J. Lezama, L. Jiang, M.-H. Yang, K. P. Murphy, W. T. Freeman, M. Rubinstein, Y. Li, and D. Krishnan · 2023
Cited alongside, same era.
Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion Models
H. Chefer, Y. Alaluf, Y. Vinker, L. Wolf, and D. Cohen-Or · 2023
Cited alongside, same era.
X. Chen, Y. Liu, Y. Yang, J. Yuan, Q. You, L.-P. Liu, and H. Yang · 2023
Cited alongside, same era.
Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models
J. Cho, A. Zala, and M. Bansal · 2023
Cited alongside, same era.
Phi-2: The surprising power of small language models
M. Javaheripi, S. Bubeck, M. Abdin, J. Aneja, S. Bubeck, C. C. T. Mendes, W. Chen, A. Del Giorno, R. Eldan, S. Gopi, et al · 2023
Cited alongside, same era.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed · 2023
Cited alongside, same era.
Ultralytics YOLO, 2023
G. Jocher, A. Chaurasia, and J. Qiu · 2023
Cited alongside, same era.
L. Zhang, A. Rao, and M. Agrawala · 2023
Later among the works it cites.
Training-free layout control with cross-attention guidance
M. Chen, I. Laina, and A. Vedaldi · 2024
Closest in time.
LLM-grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models
L. Lian, B. Li, A. Yala, and T. Darrell · 2024
Closest in time.
Directed diffusion: Direct control of object placement through attention guidance
W.-D. K. Ma, A. Lahiri, J. P. Lewis, T. Leung, and W. B. Kleijn · 2024
Closest in time.
Gpt-4 technical report, 2024
OpenAI · 2024
Closest in time.
Grounded Text-to-Image Synthesis with Attention Refocusing
Q. Phung, S. Ge, and J.-B. Huang · 2024
Closest in time.
Tinyllama: An open-source small language model
P. Zhang, G. Zeng, T. Wang, and W. Lu · 2024
Closest in time.
Qwen3, April 2025
Q. Team · 2025
Closest in time.
Fastcomposer: Tuning-free multi-subject image generation with localized attention
G. Xiao, T. Yin, W. T. Freeman, F. Durand, and S. Han · 2025
Closest in time.