Fetching the paper…
Reading the bibliography…
Recent text-to-image systems face limitations in handling multimodal inputs and complex reasoning tasks.
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 10684–10695 (2022)
2022
Earlier work this paper cites.
Schuhmann, C., Köpf, A., Vencu, R., Coombes, T., Beaumont, R.: Laion coco: 600m synthetic captions from laion2b-en. URL https://laion. ai/blog/laion-coco 5
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q.V., Zhou, D., et al.: Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Ghosh, D., Hajishirzi, H., Schmidt, L.: Geneval: An object-focused framework for evaluating text-to-image alignment. Advances in Neural Information Processing Systems pp. 52132–52152 (2023)
2023
Earlier work this paper cites.
Labs, B.F.: Flux. https://github.com/black-forest-labs/flux (2023)
2023
Earlier work this paper cites.
Peebles, W., Xie, S.: Scalable diffusion models with transformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 4195–4205 (2023)
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Chen, Z., Wang, W., Tian, H., Ye, S., Gao, Z., Cui, E., Tong, W., Hu, K., Luo, J., Ma, Z., et al.: How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites. Science China Information Sciences 67
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Cited alongside, same era.
OpenAI: Openai o1. https://openai.com/o1 (2024)
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Yang, L., Yu, Z., Meng, C., Xu, M., Ermon, S., Cui, B.: Mastering text-to-image diffusion: Recaptioning, planning, and generating with multimodal llms. In: Forty-first International Conference on Machine Learning (2024)
2024
Later among the works it cites.
Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y., et al.: Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 9556–9567 (2024)
2024
Later among the works it cites.
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
stability.ai: Introducing stable diffusion 3.5 (2024), https://stability.ai/news/introducing-stable-diffusion-3-5
2024
Cited alongside, same era.
Sun, K., Pan, J., Ge, Y., Li, H., Duan, H., Wu, X., Zhang, R., Zhou, A., Qin, Z., Wang, Y., et al.: Journeydb: A benchmark for generative image understanding. Advances in Neural Information Processing Systems 36
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Sun, Q., Cui, Y., Zhang, X., Zhang, F., Yu, Q., Wang, Y., Rao, Y., Liu, J., Huang, T., Wang, X.: Generative multimodal models are in-context learners. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 14398–14409 (2024)
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2025
Closest in time.
Chen, X., Wu, Z., Liu, X., Pan, Z., Liu, W., Xie, Z., Yu, X., Ruan, C.: Janus-pro: Unified multimodal understanding and generation with data and model scaling (2025)
2025
Closest in time.
2025
Closest in time.
Google: Gemini 2.5: Our most intelligent ai model (2025), https://blog.google/technology/google-deepmind/
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
OpenAI: Introducing 4o image generation (2025), https://openai.com
2025
Closest in time.
2025
Closest in time.
Team, Q.: Qwen3: Think deeper, act faster. https://qwenlm.github.io/blog/qwen3/ (2025)
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.