Fetching the paper…
Reading the bibliography…
The rapid advancements of Text-to-Image (T2I) models have ushered in a new phase of AI-generated content, marked by their growing ability to interpret and follow user instructions.
2021
Earlier work this paper cites.
Hessel, J., Holtzman, A., Forbes, M., Bras, R.L., Choi, Y.: CLIPScore: A Reference-free Evaluation Metric for Image Captioning (Mar 2022). https://doi.org/10.48550/arXiv.2104.08718
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-Resolution Image Synthesis With Latent Diffusion Models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 10684–10695 (2022)
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Betker, J., Goh, G., Jing, L., Brooks, T., Wang, J., Li, L., Ouyang, L., Zhuang, J., Lee, J., Guo, Y., et al.: Improving image generation with better captions. Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf 2
2023
Earlier work this paper cites.
Chen, J., Yu, J., Ge, C., Yao, L., Xie, E., Wu, Y., Wang, Z., Kwok, J., Luo, P., Lu, H., Li, Z.: PixArt-$ α $ \alpha\$ : Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis (Dec 2023). https://doi.org/10.48550/arXiv.2310.00426
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., Rombach, R.: SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis (Jul 2023). https://doi.org/10.48550/arXiv.2307.01952
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2024
Earlier work this paper cites.
Chen, J., Ge, C., Xie, E., Wu, Y., Yao, L., Ren, X., Wang, Z., Luo, P., Lu, H., Li, Z.: PixArt- Σ \Sigma : Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation (Mar 2024). https://doi.org/10.48550/arXiv.2403.04692
2024
Earlier work this paper cites.
Chen, J., Wu, Y., Luo, S., Xie, E., Paul, S., Luo, P., Zhao, H., Li, Z.: PIXART- δ \delta : Fast and Controllable Image Generation with Latent Consistency Models (Jan 2024). https://doi.org/10.48550/arXiv.2401.05252
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
Cui, S., Guo, J., An, X., Deng, J., Zhao, Y., Wei, X., Feng, Z.: IDAdapter: Learning mixed features for tuning-free personalization of text-to-image models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops. pp. 950–959 (June 2024)
2024
Earlier work this paper cites.
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., Podell, D., Dockhorn, T., English, Z., Rombach, R.: Scaling Rectified Flow Transformers for High-Resolution Image Synthesis. In: Forty-First International Conference on Machine Learning (Jun 2024)
2024
Earlier work this paper cites.
Han, J., Liu, J., Jiang, Y., Yan, B., Zhang, Y., Yuan, Z., Peng, B., Liu, X.: Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis (2024)
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
Imagen-Team-Google, Baldridge, J., Bauer, J., Bhutani, M., Others: Imagen 3 (Dec 2024). https://doi.org/10.48550/arXiv.2408.07009
2024
Earlier work this paper cites.
Labs, B.F.: Flux. https://github.com/black-forest-labs/flux (2024)
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
Li, Z., Zhang, J., Lin, Q., Xiong, J., Long, Y., Deng, X., Zhang, Y., Liu, X., Huang, M., Xiao, Z., Chen, D., He, J., Li, J., Li, W., Zhang, C., Quan, R., Lu, J., Huang, J., Yuan, X., Zheng, X., Li, Y., Zhang, J., Zhang, C., Chen, M., Liu, J., Fang, Z., Wang, W., Xue, J., Tao, Y., Zhu, J., Liu, K., Lin, S., Sun, Y., Li, Y., Wang, D., Chen, M., Hu, Z., Xiao, X., Chen, Y., Liu, Y., Liu, W., Wang, D., Yang, Y., Jiang, J., Lu, Q.: Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding (May 2024). https://doi.org/10.48550/arXiv.2405.08748
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
Sun, P., Jiang, Y., Chen, S., Zhang, S., Peng, B., Luo, P., Yuan, Z.: Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation (Jun 2024). https://doi.org/10.48550/arXiv.2406.06525
2024
Cited alongside, same era.
Tian, K., Jiang, Y., Yuan, Z., Peng, B., Wang, L.: Visual autoregressive modeling: Scalable image generation via next-scale prediction (2024)
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Wu, C., Chen, X., Wu, Z., Ma, Y., Liu, X., Pan, Z., Liu, W., Xie, Z., Yu, X., Ruan, C., Luo, P.: Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation (Oct 2024). https://doi.org/10.48550/arXiv.2410.13848
2024
Cited alongside, same era.
2025
Closest in time.
2025
Closest in time.
Team, M.: Midjourney. https://www.midjourney.com/ (2025)
2025
Closest in time.
Tong, C., Guo, Z., Zhang, R., Shan, W., Wei, X., Xing, Z., Li, H., Heng, P.A.: Delving into RL for image generation with CoT: A study on DPO vs. GRPO. In: Advances in Neural Information Processing Systems (2025)
2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
Xie, J., Mao, W., Bai, Z., Zhang, D.J., Wang, W., Lin, K.Q., Gu, Y., Chen, Z., Yang, Z., Shou, M.Z.: Show-o: One Single Transformer to Unify Multimodal Understanding and Generation (Oct 2024). https://doi.org/10.48550/arXiv.2408.12528
2024
Cited alongside, same era.
Zhuo, L., Du, R., Xiao, H., Li, Y., Liu, D., Huang, R., Liu, W., Zhao, L., Wang, F.Y., Ma, Z., Luo, X., Wang, Z., Zhang, K., Zhu, X., Liu, S., Yue, X., Liu, D., Ouyang, W., Liu, Z., Qiao, Y., Li, H., Gao, P.: Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT (Jun 2024). https://doi.org/10.48550/arXiv.2406.18583
2024
Cited alongside, same era.
2025
Cited alongside, same era.
2025
Cited alongside, same era.
Chen, J., Xue, S., Zhao, Y., Yu, J., Paul, S., Chen, J., Cai, H., Xie, E., Han, S.: SANA-Sprint: One-Step Diffusion with Continuous-Time Consistency Distillation (Mar 2025). https://doi.org/10.48550/arXiv.2503.09641
2025
Cited alongside, same era.
2025
Cited alongside, same era.
2025
Cited alongside, same era.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Xie, E., Chen, J., Zhao, Y., Yu, J., Zhu, L., Wu, C., Lin, Y., Zhang, Z., Li, M., Chen, J., Cai, H., Liu, B., Zhou, D., Han, S.: SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer (Mar 2025). https://doi.org/10.48550/arXiv.2501.18427
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Zhang, R., Wei, X., Jiang, D., Guo, Z., Li, S., Zhang, Y., Tong, C., Liu, J., Zhou, A., Wei, B., Zhang, S., Gao, P., Li, C., Li, H.: MAVIS: Mathematical visual instruction tuning with an automatic data engine. In: International Conference on Learning Representations (2025)
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2026
Closest in time.
Google: nano-banana-2. https://deepmind.google/models/gemini-image/flash/ (2026)
2026
Closest in time.
Labs, B.F.: Flux.2. https://bfl.ai/models/flux-2 (2026)
2026
Closest in time.
2026
Closest in time.
2026
Closest in time.
2026
Closest in time.
Wei, X., Cen, K., Wei, H., Guo, Z., Li, B., Wang, Z., Zhang, J., Zhang, L.: MICo-150K: A comprehensive dataset advancing multi-image composition. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 29695–29706 (June 2026)
2026
Closest in time.