Fetching the paper…
Reading the bibliography…
Accurately controlling object count in text-to-image generation remains a key challenge.
Number sense in infancy predicts mathematical abilities in childhood
Ariel Starr, Melissa E Libertus, and Elizabeth M Brannon. 2013 · 2013
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Learning to count objects in natural images for visual question answering
Yan Zhang, Jonathon Hare, and Adam Prügel-Bennett. 2018 · 2018
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol. 2021 · 2021
Earlier work this paper cites.
Clipscore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021 · 2021
Earlier work this paper cites.
Classifier-Free Diffusion Guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications
Jonathan Ho and Tim Salimans. 2021 · 2021
Earlier work this paper cites.
Diffusionclip: Text-guided image manipulation using diffusion models
Gwanghyun Kim and Jong Chul Ye. 2021 · 2021
Earlier work this paper cites.
Learning Transferable Visual Models From Natural Language Supervision. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event (Proceedings of Machine Learning Research, Vol. 139) , Marina Meila and Tong Zhang (Eds.). PMLR, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Earlier work this paper cites.
Learning to count everything. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 3394–3403
Viresh Ranjan, Udbhav Sharma, Thu Nguyen, and Minh Hoai. 2021 · 2021
Earlier work this paper cites.
VQGAN-CLIP: Open Domain Image Generation and Editing with Natural Language Guidance. In Computer Vision – ECCV 2022 , Shai Avidan, Gabriel Brostow, Moustapha Cissé, Giovanni Maria Farinella, and Tal Hassner (Eds.). Springer Nature Switzerland, Cham, 88–105
Katherine Crowson, Stella Biderman, Daniel Kornis, Dashiell Stander, Eric Hallahan, Louis Castricato, and Edward Raff. 2022 · 2022
Earlier work this paper cites.
CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers
Ming Ding, Wendi Zheng, Wenyi Hong, and Jie Tang. 2022 · 2022
Earlier work this paper cites.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022 · 2022
Earlier work this paper cites.
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding. In Advances in Neural Information Processing Systems , Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho (Eds.)
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo-Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. 2022 · 2022
Cited alongside, same era.
Scaling Autoregressive Models for Content-Rich Text-to-Image Generation
Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, Ben Hutchinson, Wei Han, Zarana Parekh, Xin Li, Han Zhang, Jason Baldridge, and Yonghui Wu. 2022 · 2022
Cited alongside, same era.
Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models
Hila Chefer, Yuval Alaluf, Yael Vinker, Lior Wolf, and Daniel Cohen-Or. 2023 · 2023
Cited alongside, same era.
An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion. In The Eleventh International Conference on Learning Representations
Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit Haim Bermano, Gal Chechik, and Daniel Cohen-or. 2023 · 2023
Cited alongside, same era.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. 2023 · 2023
Later among the works it cites.
Adversarial diffusion distillation
Axel Sauer, Dominik Lorenz, Andreas Blattmann, and Robin Rombach. 2023 · 2023
Later among the works it cites.
Discriminative class tokens for text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 22725–22735
Idan Schwartz, Vésteinn Snæbjarnarson, Hila Chefer, Serge Belongie, Lior Wolf, and Sagie Benaim. 2023 · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 3836–3847
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Prompt-to-Prompt Image Editing with Cross-Attention Control. In The Eleventh International Conference on Learning Representations
Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-or. 2023 · 2023
Cited alongside, same era.
YOLO-v1 to YOLO-v8, the rise of YOLO and its complementary nature toward digital manufacturing and industrial defect detection
Muhammad Hussain. 2023 · 2023
Cited alongside, same era.
CLIP-Count: Towards Text-Guided Zero-Shot Object Counting
Ruixiang Jiang, Lingbo Liu, and Changwen Chen. 2023 · 2023
Cited alongside, same era.
GLIGEN: Open-Set Grounded Text-to-Image Generation
Yuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu, Jianwei Yang, Jianfeng Gao, Chunyuan Li, and Yong Jae Lee. 2023 · 2023
Cited alongside, same era.
Zero-1-to-3: Zero-shot One Image to 3D Object
Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. 2023 · 2023
Cited alongside, same era.
Null-text inversion for editing real images using guided diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 6038–6047
Ron Mokady, Amir Hertz, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. 2023 · 2023
Cited alongside, same era.
Teaching clip to count to ten. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 3170–3180
Roni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada, Inbar Mosseri, Michal Irani, and Tali Dekel. 2023 · 2023
Cited alongside, same era.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2023 · 2023
Cited alongside, same era.
Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion Models
Shihao Zhao, Dongdong Chen, Yen-Chun Chen, Jianmin Bao, Shaozhe Hao, Lu Yuan, and Kwan-Yee K Wong. 2023 · 2023
Later among the works it cites.
Make It Count: Text-to-Image Generation with an Accurate Number of Objects
Lital Binyamin, Yoad Tewel, Hilit Segev, Eran Hirsch, Royi Rassin, and Gal Chechik. 2024 · 2024
Closest in time.
Training Diffusion Models with Reinforcement Learning. In International Conference on Learning Representations (ICLR)
Jacob Austin Black, Sitao Xie, Yilun Du, William T. Freeman, Joshua B. Tenenbaum, Andrew Zisserman, and Roger Grosse. 2024 · 2024
Closest in time.
AFreeCA: Annotation-Free Counting for All
Adriano D’Alessandro, Ali Mahdavi-Amiri, and Ghassan Hamarneh. 2024 · 2024
Closest in time.
Initno: Boosting text-to-image diffusion models via initial noise optimization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 9380–9389
Xiefan Guo, Jinlin Liu, Miaomiao Cui, Jiankai Li, Hongyu Yang, and Di Huang. 2024 · 2024
Closest in time.
Instaflow: One step is enough for high-quality diffusion-based text-to-image generation. In International Conference on Learning Representations
Xingchao Liu, Xiwen Zhang, Jianzhu Ma, Jian Peng, and Qiang Liu. 2024 · 2024
Closest in time.
Generating images of rare concepts using pre-trained diffusion models. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 4695–4703
Dvir Samuel, Rami Ben-Ari, Simon Raviv, Nir Darshan, and Gal Chechik. 2024 · 2024
Closest in time.
Counting Guidance for High Fidelity Text-to-Image Synthesis. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) . IEEE, 899–908
Wonjun Kang, Kevin Galim, Hyung Il Koo, and Nam Ik Cho. 2025 · 2025
Closest in time.