Fetching the paper…
Reading the bibliography…
In this work, we propose a training-free, trajectory-based controllable T2I approach, termed TraDiffusion.
The opencv library
Bradski, G. 2000 · 2000
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J.; Meng, C.; and Ermon, S. 2020 · 2010
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Song, Y.; Sohl-Dickstein, J.; Kingma, D. P.; Kumar, A.; Ermon, S.; and Poole, B. 2020 · 2011
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P.; and Welling, M. 2013 · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Zeiler, M. D.; and Fergus, R. 2014 · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O.; Fischer, P.; and Brox, T. 2015 · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J.; Weiss, E.; Maheswaranathan, N.; and Ganguli, S. 2015 · 2015
Earlier work this paper cites.
You only look once: Unified, real-time object detection
Redmon, J.; Divvala, S.; Girshick, R.; and Farhadi, A. 2016 · 2016
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks
Isola, P.; Zhu, J.-Y.; Zhou, T.; and Efros, A. A. 2017 · 2017
Earlier work this paper cites.
Unsupervised image-to-image translation networks
Liu, M.-Y.; Breuel, T.; and Kautz, J. 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A.; Vinyals, O.; et al. 2017 · 2017
Earlier work this paper cites.
Unpaired image-to-image translation using cycle-consistent adversarial networks
Zhu, J.-Y.; Park, T.; Isola, P.; and Efros, A. A. 2017 · 2017
Earlier work this paper cites.
Squeeze-and-excitation networks
Hu, J.; Shen, L.; and Sun, G. 2018 · 2018
Earlier work this paper cites.
Image generation from scene graphs
Johnson, J.; Gupta, A.; and Fei-Fei, L. 2018 · 2018
Earlier work this paper cites.
Attention u-net: Learning where to look for the pancreas
Oktay, O.; Schlemper, J.; Folgoc, L. L.; Lee, M.; Heinrich, M.; Misawa, K.; Mori, K.; McDonagh, S.; Hammerla, N. Y.; Kainz, B.; et al. 2018 · 2018
Earlier work this paper cites.
High-resolution image synthesis and semantic manipulation with conditional gans
Wang, T.-C.; Liu, M.-Y.; Zhu, J.-Y.; Tao, A.; Kautz, J.; and Catanzaro, B. 2018 · 2018
Earlier work this paper cites.
Attngan: Fine-grained text to image generation with attentional generative adversarial networks
Xu, T.; Zhang, P.; Huang, Q.; Zhang, H.; Gan, Z.; Huang, X.; and He, X. 2018 · 2018
Earlier work this paper cites.
Semantic image synthesis with spatially-adaptive normalization
Park, T.; Liu, M.-Y.; Wang, T.-C.; and Zhu, J.-Y. 2019 · 2019
Earlier work this paper cites.
Image synthesis from reconfigurable layout and style
Sun, W.; and Wu, T. 2019 · 2019
Earlier work this paper cites.
Self-attention generative adversarial networks
Zhang, H.; Goodfellow, I.; Metaxas, D.; and Odena, A. 2019 · 2019
Earlier work this paper cites.
Image generation from layout
Zhao, B.; Meng, L.; Yin, W.; and Sigal, L. 2019 · 2019
Earlier work this paper cites.
Generative adversarial networks
Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y. 2020 · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Cited alongside, same era.
Bachgan: High-resolution image synthesis from salient object layout
Li, Y.; Cheng, Y.; Gan, Z.; Yu, L.; Wang, L.; and Liu, J. 2020 · 2020
Cited alongside, same era.
Connecting vision and language with localized narratives
Pont-Tuset, J.; Uijlings, J.; Changpinyo, S.; Soricut, R.; and Ferrari, V. 2020 · 2020
Cited alongside, same era.
Text-guided neural image inpainting
Zhang, L.; Chen, Q.; Hu, B.; and Jiang, S. 2020 · 2020
Cited alongside, same era.
Diffusion models beat gans on image synthesis
Dhariwal, P.; and Nichol, A. 2021 · 2021
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E. L.; Ghasemipour, K.; Gontijo Lopes, R.; Karagol Ayan, B.; Salimans, T.; et al. 2022 · 2022
Later among the works it cites.
Cross-view panorama image synthesis
Wu, S.; Tang, H.; Jing, X.-Y.; Zhao, H.; Qian, J.; Sebe, N.; and Yan, Y. 2022 · 2022
Later among the works it cites.
Modeling image composition for complex scene generation
Yang, Z.; Liu, D.; Wang, C.; Yang, J.; and Tao, D. 2022 · 2022
Later among the works it cites.
Spatext: Spatio-textual representation for controllable image generation
Avrahami, O.; Hayes, T.; Gafni, O.; Gupta, S.; Taigman, Y.; Parikh, D.; Lischinski, D.; Fried, O.; and Yin, X. 2023 · 2023
Later among the works it cites.
Multidiffusion: Fusing diffusion paths for controlled image generation
Bar-Tal, O.; Yariv, L.; Lipman, Y.; and Dekel, T. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Harmonious textual layout generation over natural images via deep aesthetics learning
Li, C.; Zhang, P.; and Wang, C. 2021 · 2021
Cited alongside, same era.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Nichol, A.; Dhariwal, P.; Ramesh, A.; Shyam, P.; Mishkin, P.; McGrew, B.; Sutskever, I.; and Chen, M. 2021 · 2021
Cited alongside, same era.
Layout Structure Assisted Indoor Image Generation
Qin, Z.; Zhong, W.; Hu, F.; Yang, X.; Ye, L.; and Zhang, Q. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Object-centric image generation from layouts
Sylvain, T.; Zhang, P.; Bengio, Y.; Hjelm, R. D.; and Sharma, S. 2021 · 2021
Cited alongside, same era.
Improving text-to-image generation with object layout guidance
Zakraoui, J.; Saleh, M.; Al-Maadeed, S.; and Jaam, J. M. 2021 · 2021
Cited alongside, same era.
Huang, L.; Chen, D.; Liu, Y.; Shen, Y.; Zhao, D.; and Zhou, J. 2023 · 2023
Later among the works it cites.
Ultralytics YOLOv8
Jocher, G.; Chaurasia, A.; and Qiu, J. 2023 · 2023
Later among the works it cites.
Dense text-to-image generation with attention modulation
Kim, Y.; Lee, J.; Kim, J.-H.; Ha, J.-W.; and Zhu, J.-Y. 2023 · 2023
Later among the works it cites.
Gligen: Open-set grounded text-to-image generation
Li, Y.; Liu, H.; Wu, Q.; Mu, F.; Yang, J.; Gao, J.; Li, C.; and Lee, Y. J. 2023 · 2023
Later among the works it cites.
Layoutllm-t2i: Eliciting layout guidance from llm for text-to-image generation
Qu, L.; Wu, S.; Fei, H.; Nie, L.; and Chua, T.-S. 2023 · 2023
Later among the works it cites.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Ruiz, N.; Li, Y.; Jampani, V.; Pritch, Y.; Rubinstein, M.; and Aberman, K. 2023 · 2023
Later among the works it cites.
Alr-gan: Adaptive layout refinement for text-to-image synthesis
Tan, H.; Yin, B.; Wei, K.; Liu, X.; and Li, X. 2023 · 2023
Later among the works it cites.
Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion
Xie, J.; Li, Y.; Huang, Y.; Liu, H.; Zhang, W.; Zheng, Y.; and Shou, M. Z. 2023 · 2023
Later among the works it cites.
Xu, J.; Zhou, X.; Yan, S.; Gu, X.; Arnab, A.; Sun, C.; Wang, X.; and Schmid, C. 2023 · 2023
Later among the works it cites.
Reco: Region-controlled text-to-image generation
Yang, Z.; Wang, J.; Gan, Z.; Li, L.; Lin, K.; Wu, C.; Duan, N.; Liu, Z.; Liu, C.; Zeng, M.; et al. 2023 · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
Zhang, L.; Rao, A.; and Agrawala, M. 2023 · 2023
Later among the works it cites.
Training-free layout control with cross-attention guidance
Chen, M.; Laina, I.; and Vedaldi, A. 2024 · 2024
Closest in time.
Layoutgpt: Compositional visual planning and generation with large language models
Feng, W.; Zhu, W.; Fu, T.-j.; Jampani, V.; Akula, A.; He, X.; Basu, S.; Wang, X. E.; and Wang, W. Y. 2024 · 2024
Closest in time.
Diffusion model-based image editing: A survey
Huang, Y.; Huang, J.; Liu, Y.; Yan, M.; Lv, J.; Liu, J.; Xiong, W.; Zhang, H.; Chen, S.; and Cao, L. 2024 · 2024
Closest in time.
Move Anything with Layered Scene Diffusion
Ren, J.; Xu, M.; Wu, J.-C.; Liu, Z.; Xiang, T.; and Toisoul, A. 2024 · 2024
Closest in time.
InstanceDiffusion: Instance-level Control for Image Generation
Wang, X.; Darrell, T.; Rambhatla, S. S.; Girdhar, R.; and Misra, I. 2024 · 2024
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K.; Ba, J.; Kiros, R.; Cho, K.; Courville, A.; Salakhudinov, R.; Zemel, R.; and Bengio, Y. 2015 · 2057
Closest in time.