Fetching the paper…
Reading the bibliography…
We introduce SLayR, Scene Layout Generation with Rectified flow, a novel transformer-based model for text-to-layout generation which can then be paired with existing layout-to-image models to produce images.
Microsoft coco: Common objects in context, 2015
Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics, 2015
Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium, 2018
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2018
Earlier work this paper cites.
Semantic understanding of scenes through the ade20k dataset, 2018
Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fidler, Adela Barriuso, and Antonio Torralba · 2018
Earlier work this paper cites.
Layoutgan: Generating graphic layouts with wireframe discriminators, 2019
Jianan Li, Jimei Yang, Aaron Hertzmann, Jianming Zhang, and Tingfa Xu · 2019
Earlier work this paper cites.
Image generation from layout, 2019
Bo Zhao, Lili Meng, Weidong Yin, and Leonid Sigal · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models, 2020
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Neural design network: Graphic layout generation with constraints, 2020
Hsin-Ying Lee, Lu Jiang, Irfan Essa, Phuong B Le, Haifeng Gong, Ming-Hsuan Yang, and Weilong Yang · 2020
Earlier work this paper cites.
Object-centric image generation from layouts, 2020
Tristan Sylvain, Pengchuan Zhang, Y. Bengio, R Devon Hjelm, and Shikhar Sharma · 2020
Earlier work this paper cites.
Demystifying mmd gans, 2021
Mikołaj Bińkowski, Danica J. Sutherland, Michael Arbel, and Arthur Gretton · 2021
Earlier work this paper cites.
Layouttransformer: Layout generation and completion with self-attention, 2021
Kamal Gupta, Justin Lazarow, Alessandro Achille, Larry Davis, Vijay Mahadevan, and Abhinav Shrivastava · 2021
Earlier work this paper cites.
Constrained graphic layout generation via latent optimization
Kotaro Kikuchi, Edgar Simo-Serra, Mayu Otani, and Kota Yamaguchi · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation, 2021
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Earlier work this paper cites.
Flow straight and fast: Learning to generate and transfer data with rectified flow, 2022
Xingchao Liu, Chengyue Gong, and Qiang Liu · 2022
Earlier work this paper cites.
Repaint: Inpainting using denoising diffusion probabilistic models, 2022
Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models, 2022
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Earlier work this paper cites.
Photorealistic text-to-image diffusion models with deep language understanding, 2022
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi · 2022
Earlier work this paper cites.
Denoising diffusion implicit models, 2022
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2022
Earlier work this paper cites.
Reco: Region-controlled text-to-image generation, 2022
Zhengyuan Yang, Jianfeng Wang, Zhe Gan, Linjie Li, Kevin Lin, Chenfei Wu, Nan Duan, Zicheng Liu, Ce Liu, Michael Zeng, and Lijuan Wang · 2022
Cited alongside, same era.
Multidiffusion: Fusing diffusion paths for controlled image generation, 2023
Omer Bar-Tal, Lior Yariv, Yaron Lipman, and Tali Dekel · 2023
Cited alongside, same era.
Layoutdm: Transformer-based diffusion model for layout generation, 2023
Shang Chai, Liansheng Zhuang, and Fengying Yan · 2023
Cited alongside, same era.
Training-free layout control with cross-attention guidance, 2023
Minghao Chen, Iro Laina, and Andrea Vedaldi · 2023
Cited alongside, same era.
Layoutgpt: Compositional visual planning and generation with large language models, 2023
Weixi Feng, Wanrong Zhu, Tsu jui Fu, Varun Jampani, Arjun Akula, Xuehai He, Sugato Basu, Xin Eric Wang, and William Yang Wang · 2023
Cited alongside, same era.
Graphdreamer: Compositional 3d scene synthesis from scene graphs, 2024
Gege Gao, Weiyang Liu, Anpei Chen, Andreas Geiger, and Bernhard Schölkopf · 2024
Closest in time.
Layoutflow: Flow matching for layout generation, 2024
Julian Jorge Andrade Guerreiro, Naoto Inoue, Kento Masui, Mayu Otani, and Hideki Nakayama · 2024
Closest in time.
Rethinking fid: Towards a better evaluation metric for image generation, 2024
Sadeep Jayasumana, Srikumar Ramalingam, Andreas Veit, Daniel Glasner, Ayan Chakrabarti, and Sanjiv Kumar · 2024
Closest in time.
Llm-grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models, 2024
Long Lian, Boyi Li, Adam Yala, and Trevor Darrell · 2024
Closest in time.
Rich human feedback for text-to-image generation, 2024
Youwei Liang, Junfeng He, Gang Li, Peizhao Li, Arseniy Klimovskiy, Nicholas Carolan, Jiao Sun, Jordi Pont-Tuset, Sarah Young, Feng Yang, Junjie Ke, Krishnamurthy Dj Dvijotham, Katie Collins, Yiwen Luo, Yang Li, Kai J Kohlhoff, Deepak Ramachandran, and Vidhya Navalpakkam · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Drew A. Hudson, Daniel Zoran, Mateusz Malinowski, Andrew K. Lampinen, Andrew Jaegle, James L. McClelland, Loic Matthey, Felix Hill, and Alexander Lerchner · 2023
Cited alongside, same era.
Layoutdm: Discrete diffusion model for controllable layout generation, 2023
Naoto Inoue, Kotaro Kikuchi, Edgar Simo-Serra, Mayu Otani, and Kota Yamaguchi · 2023
Cited alongside, same era.
Layoutformer++: Conditional graphic layout generation via constraint serialization and decoding space restriction
Zhaoyun Jiang, Jiaqi Guo, Shizhao Sun, Huayu Deng, Zhongkai Wu, Vuksan Mijovic, Zijiang James Yang, Jian-Guang Lou, and Dongmei Zhang · 2023
Cited alongside, same era.
Diffusion models already have a semantic latent space, 2023
Mingi Kwon, Jaeseok Jeong, and Youngjung Uh · 2023
Cited alongside, same era.
Dlt: Conditioned layout generation with joint discrete-continuous diffusion layout transformer, 2023
Elad Levi, Eli Brosh, Mykola Mykhailych, and Meir Perez · 2023
Cited alongside, same era.
Gligen: Open-set grounded text-to-image generation, 2023
Yuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu, Jianwei Yang, Jianfeng Gao, Chunyuan Li, and Yong Jae Lee · 2023
Cited alongside, same era.
Toward verifiable and reproducible human evaluation for text-to-image generation, 2023
Mayu Otani, Riku Togashi, Yu Sawai, Ryosuke Ishigami, Yuta Nakashima, Esa Rahtu, Janne Heikkilä, and Shin’ichi Satoh · 2023
Cited alongside, same era.
Evaluating text-to-visual generation with image-to-text generation, 2024
Zhiqiu Lin, Deepak Pathak, Baiqi Li, Jiayao Li, Xide Xia, Graham Neubig, Pengchuan Zhang, and Deva Ramanan · 2024
Closest in time.
Instaflow: One step is enough for high-quality diffusion-based text-to-image generation, 2024
Xingchao Liu, Xiwen Zhang, Jianzhu Ma, Jian Peng, and Qiang Liu · 2024
Closest in time.
Readout guidance: Learning control from diffusion features, 2024
Grace Luo, Trevor Darrell, Oliver Wang, Dan B Goldman, and Aleksander Holynski · 2024
Closest in time.
Image synthesis with graph conditioning: Clip-guided diffusion models for scene graphs, 2024
Rameshwar Mishra and A V Subramanyam · 2024
Closest in time.
Gpt-4 technical report, 2024
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, et al · 2024
Closest in time.
Fast high-resolution image synthesis with latent adversarial diffusion distillation, 2024
Axel Sauer, Frederic Boesel, Tim Dockhorn, Andreas Blattmann, Patrick Esser, and Robin Rombach · 2024
Closest in time.
Neural point cloud diffusion for disentangled 3d shape and appearance generation, 2024
Philipp Schröppel, Christopher Wewer, Jan Eric Lenssen, Eddy Ilg, and Thomas Brox · 2024
Closest in time.
Sg-adapter: Enhancing text-to-image generation with scene graph guidance, 2024
Guibao Shen, Luozhou Wang, Jiantao Lin, Wenhang Ge, Chaozhe Zhang, Xin Tao, Yuan Zhang, Pengfei Wan, Zhongyuan Wang, Guangyong Chen, Yijun Li, and Ying-Cong Chen · 2024
Closest in time.
Instancediffusion: Instance-level control for image generation, 2024
Xudong Wang, Trevor Darrell, Sai Saketh Rambhatla, Rohit Girdhar, and Ishan Misra · 2024
Closest in time.
Groundingbooth: Grounding text-to-image customization, 2024
Zhexiao Xiong, Wei Xiong, Jing Shi, He Zhang, Yizhi Song, and Nathan Jacobs · 2024
Closest in time.
Dreamscape: 3d scene creation via gaussian splatting joint correlation modeling, 2024
Xuening Yuan, Hongyu Yang, Yueming Zhao, and Di Huang · 2024
Closest in time.
Gala3d: Towards text-to-3d complex scene generation via layout-guided generative gaussian splatting, 2024
Xiaoyu Zhou, Xingjian Ran, Yajiao Xiong, Jinlin He, Zhiwei Lin, Yongtao Wang, Deqing Sun, and Ming-Hsuan Yang · 2024
Closest in time.
Sceneteller: Language-to-3d scene generation, 2024
Başak Melis Öcal, Maxim Tatarchenko, Sezer Karaoglu, and Theo Gevers · 2024
Closest in time.