Fetching the paper…
Reading the bibliography…
Despite the unprecedented success of text-to-image diffusion models, controlling the number of depicted objects using text is surprisingly hard.
The hungarian method for the assignment problem
H. W. Kuhn · 1955
Earlier work this paper cites.
A threshold selection method from gray-level histograms
Nobuyuki Otsu · 1979
Earlier work this paper cites.
Principles of object perception
Elizabeth S Spelke · 1990
Earlier work this paper cites.
A density-based algorithm for discovering clusters in large spatial databases with noise
Martin Ester, Hans-Peter Kriegel, Jörg Sander, and Xiaowei Xu · 1996
Earlier work this paper cites.
Measuring the objectness of image windows
Bogdan Alexe, Thomas Deselaers, and Vittorio Ferrari · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Deepbox: Learning objectness with convolutional networks
Weicheng Kuo, Bharath Hariharan, and Jitendra Malik · 2015
Earlier work this paper cites.
Generalised dice overlap as a deep learning loss function for highly unbalanced segmentations
Carole H Sudre, Wenqi Li, Tom Vercauteren, Sebastien Ourselin, and M Jorge Cardoso · 2017
Earlier work this paper cites.
Association of genomic subtypes of lower-grade gliomas with shape features automatically extracted by a deep learning algorithm
Mateusz Buda, Ashirbani Saha, and Maciej A Mazurowski · 2019
Earlier work this paper cites.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
DALLE-2 is seeing double: Flaws in word-to-concept mapping in Text2Image models
Royi Rassin, Shauli Ravfogel, and Yoav Goldberg · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Earlier work this paper cites.
Reason out your layout: Evoking the layout master from large language models for text-to-image synthesis, 2023
Xiaohui Chen, Yongfei Liu, Yingxiang Yang, Jianbo Yuan, Quanzeng You, Li-Ping Liu, and Hongxia Yang · 2023
Cited alongside, same era.
LayoutGPT: Compositional visual planning and generation with large language models
Weixi Feng, Wanrong Zhu, Tsu-Jui Fu, Varun Jampani, Arjun Reddy Akula, Xuehai He, S Basu, Xin Eric Wang, and William Yang Wang · 2023
Cited alongside, same era.
Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models
Hila Chefer, Yuval Alaluf, Yael Vinker, Lior Wolf, and Daniel Cohen-Or · 2023
Cited alongside, same era.
Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach · 2023
Cited alongside, same era.
T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation, 2023
Kaiyi Huang, Kaiyue Sun, Enze Xie, Zhenguo Li, and Xihui Liu · 2023
Grounded text-to-image synthesis with attention refocusing, 2023
Quynh Phung, Songwei Ge, and Jia-Bin Huang · 2023
Later among the works it cites.
Reco: Region-controlled text-to-image generation
Zhengyuan Yang, Jianfeng Wang, Zhe Gan, Linjie Li, Kevin Lin, Chenfei Wu, Nan Duan, Zicheng Liu, Ce Liu, Michael Zeng, and Lijuan Wang · 2023
Later among the works it cites.
Prompt-to-prompt image editing with cross-attention control
Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-or · 2023
Later among the works it cites.
Key-locked rank one editing for text-to-image personalization, 2023
Yoad Tewel, Rinon Gal, Gal Chechik, and Yuval Atzmon · 2023
Later among the works it cites.
Be yourself: Bounded attention for multi-subject text-to-image generation
Omer Dahary, Or Patashnik, Kfir Aberman, and Daniel Cohen-Or · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Improving image generation with better captions
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al · 2023
Cited alongside, same era.
Counting guidance for high fidelity text-to-image synthesis, 2023
Wonjun Kang, Kevin Galim, and Hyung Il Koo · 2023
Cited alongside, same era.
Zero-shot improvement of object counting with CLIP
Ruisu Zhang, Yicong Chen, and Kangwook Lee · 2023
Cited alongside, same era.
Teaching clip to count to ten
Roni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada, Inbar Mosseri, Michal Irani, and Tali Dekel · 2023
Cited alongside, same era.
Aligning text-to-image models using human feedback, 2023
Kimin Lee, Hao Liu, Moonkyung Ryu, Olivia Watkins, Yuqing Du, Craig Boutilier, Pieter Abbeel, Mohammad Ghavamzadeh, and Shixiang Shane Gu · 2023
Cited alongside, same era.
Reinforcement learning for fine-tuning text-to-image diffusion models
Ying Fan, Olivia Watkins, Yuqing Du, Hao Liu, Moonkyung Ryu, Craig Boutilier, Pieter Abbeel, Mohammad Ghavamzadeh, Kangwook Lee, and Kimin Lee · 2023
Cited alongside, same era.
Dreamsync: Aligning text-to-image generation with image understanding feedback, 2023
Jiao Sun, Deqing Fu, Yushi Hu, Su Wang, Royi Rassin, Da-Cheng Juan, Dana Alon, Charles Herrmann, Sjoerd van Steenkiste, Ranjay Krishna, and Cyrus Rashtchian · 2023
Cited alongside, same era.
Be yourself: Bounded attention for multi-subject text-to-image generation, 2024
Omer Dahary, Or Patashnik, Kfir Aberman, and Daniel Cohen-Or · 2024
Closest in time.
Yolov9: Learning what you want to learn using programmable gradient information
Chien-Yao Wang, I-Hau Yeh, and Hong-Yuan Mark Liao · 2024
Closest in time.
Improving compositional text-to-image generation with large vision-language models, 2024
Song Wen, Guian Fang, Renrui Zhang, Peng Gao, Hao Dong, and Dimitris N. Metaxas · 2024
Closest in time.
Obtaining favorable layouts for multiple object generation, 2024
Barak Battash, Amit Rozner, Lior Wolf, and Ofir Lindenbaum · 2024
Closest in time.
Linguistic binding in diffusion models: Enhancing attribute correspondence through attention map alignment
Royi Rassin, Eran Hirsch, Daniel Glickman, Shauli Ravfogel, Yoav Goldberg, and Gal Chechik · 2024
Closest in time.
LLM blueprint: Enabling text-to-image generation with complex and detailed prompts
Hanan Gani, Shariq Farooq Bhat, Muzammal Naseer, Salman Khan, and Peter Wonka · 2024
Closest in time.
Training-free consistent text-to-image generation
Yoad Tewel, Omri Kaduri, Rinon Gal, Yoni Kasten, Lior Wolf, Gal Chechik, and Yuval Atzmon · 2024
Closest in time.