Fetching the paper…
Reading the bibliography…
Text-to-Image (T2I) generation models have advanced rapidly in recent years, but accurately capturing spatial relationships like "above" or "to the right of" poses a persistent challenge.
Mosaics of scenes with moving objects
James Davis · 1998
Earlier work this paper cites.
Image quilting for texture synthesis and transfer
Alexei A. Efros and William T. Freeman · 2001
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam M. Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2019
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Earlier work this paper cites.
Masked-attention mask transformer for universal image segmentation
Bowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov, and Rohit Girdhar · 2022
Earlier work this paper cites.
Benchmarking spatial relationships in text-to-image generation
Tejas Gokhale, Hamid Palangi, Besmira Nushi, Vibhav Vineet, Eric Horvitz, Ece Kamar, Chitta Baral, and Yezhou Yang · 2022
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Earlier work this paper cites.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman · 2022
Earlier work this paper cites.
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Lit, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Raphael Gontijo-Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi · 2022
Earlier work this paper cites.
Laion-aesthetics
Christoph Schuhmann and Romain Beaumont · 2022
Earlier work this paper cites.
Building normalizing flows with stochastic interpolants
Michael S. Albergo and Eric Vanden-Eijnden · 2023
Earlier work this paper cites.
Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models
Eslam Mohamed Bakr, Pengzhan Sun, Xiaoqian Shen, Faizan Farooq Khan, Li Erran Li, and Mohamed Elhoseiny · 2023
Earlier work this paper cites.
Controlstyle: Text-driven stylized image generation using diffusion priors
Jingwen Chen, Yingwei Pan, Ting Yao, and Tao Mei · 2023
Earlier work this paper cites.
Scenegenie: Scene graph guided diffusion models for image synthesis
Azade Farshad, Yousef Yeganeh, Yu Chi, Chengzhi Shen, Böjrn Ommer, and Nassir Navab · 2023
Earlier work this paper cites.
Layoutgpt: Compositional visual planning and generation with large language models
Weixi Feng, Wanrong Zhu, Tsu-jui Fu, Varun Jampani, Arjun Akula, Xuehai He, Sugato Basu, Xin Eric Wang, and William Yang Wang · 2023
Earlier work this paper cites.
Geneval: An object-focused framework for evaluating text-to-image alignment
Dhruba Ghosh, Hannaneh Hajishirzi, and Ludwig Schmidt · 2023
Earlier work this paper cites.
T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation
Kaiyi Huang, Kaiyue Sun, Enze Xie, Zhenguo Li, and Xihui Liu · 2023
Earlier work this paper cites.
Segment anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al · 2023
Earlier work this paper cites.
Gligen: Open-set grounded text-to-image generation
Yuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu, Jianwei Yang, Jianfeng Gao, Chunyuan Li, and Yong Jae Lee · 2023
Earlier work this paper cites.
Flow matching for generative modeling
Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le · 2023
Cited alongside, same era.
Dall-e 3
OpenAI · 2023
Cited alongside, same era.
DALL·E 3
OpenAI · 2023
Cited alongside, same era.
Gpt-4 technical report, 2023
OpenAI · 2023
Cited alongside, same era.
Scalable diffusion models with transformers
William Peebles and Saining Xie · 2023
Cited alongside, same era.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, A. Blattmann, Tim Dockhorn, Jonas Muller, Joe Penna, and Robin Rombach · 2023
Cited alongside, same era.
Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion
Mastering text-to-image diffusion: Recaptioning, planning, and generating with multimodal llms
Ling Yang, Zhaochen Yu, Chenlin Meng, Minkai Xu, Stefano Ermon, and Bin Cui · 2024
Later among the works it cites.
Creatilayout: Siamese multimodal diffusion transformer for creative layout-to-image generation
Hui Zhang, Dexiang Hong, Yitong Wang, Jie Shao, Xinglong Wu, Zuxuan Wu, and Yu-Gang Jiang · 2024
Later among the works it cites.
Generative artificial intelligence, human creativity, and art
Eric Zhou and Dokyun Lee · 2024
Later among the works it cites.
Storydiffusion: Consistent self-attention for long-range image and video generation
Yupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng, and Qibin Hou · 2024
Later among the works it cites.
Sub: Benchmarking cbm generalization via synthetic attribute substitutions
Jessica Bader, Leander Girrbach, Stephan Alaniz, and Zeynep Akata · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jinheng Xie, Yuexiang Li, Yawen Huang, Haozhe Liu, Wentian Zhang, Yefeng Zheng, and Mike Zheng Shou · 2023
Cited alongside, same era.
Training-free layout control with cross-attention guidance
Minghao Chen, Iro Laina, and Andrea Vedaldi · 2024
Cited alongside, same era.
Scaling rectified flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al · 2024
Cited alongside, same era.
Reno: Enhancing one-step text-to-image models through reward-based noise optimization
Luca Vincent Eyring, Shyamgopal Karthik, Karsten Roth, Alexey Dosovitskiy, and Zeynep Akata · 2024
Cited alongside, same era.
Ella: Equip diffusion models with llm for enhanced semantic alignment
Xiwei Hu, Rui Wang, Yixiao Fang, Bin Fu, Pei Cheng, and Gang Yu · 2024
Cited alongside, same era.
Datadream: Few-shot guided dataset generation
Jae Myung Kim, Jessica Bader, Stephan Alaniz, Cordelia Schmid, and Zeynep Akata · 2024
Cited alongside, same era.
Advancing vision-language models for open-vocabulary recognition and generative evaluation
Maria A. Bravo · 2025
Closest in time.
Hidream-i1: A high-efficient image generative foundation model with sparse diffusion transformer
Qi Cai, Jingwen Chen, Yang Chen, Yehao Li, Fuchen Long, Yingwei Pan, Zhaofan Qiu, Yiheng Zhang, Fengbin Gao, Peihan Xu, et al · 2025
Closest in time.
Imagen 4
Google DeepMind · 2025
Closest in time.
Emerging properties in unified multimodal pretraining
Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou, Feng Li, Zeyu Wang, Shu Zhong, Weihao Yu, Xiaonan Nie, Ziang Song, et al · 2025
Closest in time.
Noise hypernetworks: Amortizing test-time compute in diffusion models
Luca Vincent Eyring, Shyamgopal Karthik, Alexey Dosovitskiy, Nataniel Ruiz, and Zeynep Akata · 2025
Closest in time.
Rongyao Fang, Chengqi Duan, Kun Wang, Linjiang Huang, Hao Li, Shilin Yan, Hao Tian, Xingyu Zeng, Rui Zhao, Jifeng Dai, et al · 2025
Closest in time.
T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-Image Generation
Kaiyi Huang, Chengqi Duan, Kaiyue Sun, Enze Xie, Zhenguo Li, and Xihui Liu · 2025
Closest in time.
Uniedit-flow: Unleashing inversion and editing in the era of flow models, 2025
Guanlong Jiao, Biqing Huang, Kuan-Chieh Wang, and Renjie Liao · 2025
Closest in time.
Scalable ranked preference optimization for text-to-image generation
Shyamgopal Karthik, Huseyin Coskun, Zeynep Akata, S. Tulyakov, Jian Ren, and Anil Kag · 2025
Closest in time.
Unieval: Unified holistic evaluation for unified multimodal understanding and generation
Yi Li, Haonan Wang, Qixiang Zhang, Boyu Xiao, Chenchang Hu, Hualiang Wang, and Xiaomeng Li · 2025
Closest in time.
Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation
Yiyang Ma, Xingchao Liu, Xiaokang Chen, Wen Liu, Chengyue Wu, Zhiyu Wu, Zizheng Pan, Zhenda Xie, Haowei Zhang, Xingkai Yu, et al · 2025
Closest in time.
midjourney v7, 2025
Midjourney · 2025
Closest in time.
Gpt-5, 2025
OpenAI · 2025
Closest in time.
Semu: Singular value decomposition for efficient machine unlearning
Marcin Sendera, Łukasz Struski, Kamil Ksiażek, Kryspin Musiol, Jacek Tabor, and Dawid Rymarczyk · 2025
Closest in time.