Fetching the paper…
Reading the bibliography…
Recent text-to-image (T2I) diffusion models show outstanding performance in generating high-quality images conditioned on textual prompts.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O.; Fischer, P.; and Brox, T. 2015 · 2015
Earlier work this paper cites.
End-to-end people detection in crowded scenes
Stewart, R.; Andriluka, M.; and Ng, A. Y. 2016 · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; and Hochreiter, S. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Generalized intersection over union: A metric and a loss for bounding box regression
Rezatofighi, H.; Tsoi, N.; Gwak, J.; Sadeghian, A.; Reid, I.; and Savarese, S. 2019 · 2019
Earlier work this paper cites.
End-to-end object detection with transformers
Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020 · 2020
Earlier work this paper cites.
Generative adversarial networks
Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y. 2020 · 2020
Earlier work this paper cites.
spaCy: Industrial-strength Natural Language Processing in Python
Honnibal, M.; Montani, I.; Van Landeghem, S.; and Boyd, A. 2020 · 2020
Earlier work this paper cites.
Tutorial on Variational Autoencoders
DOERSCH, C. 2021 · 2021
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Esser, P.; Rombach, R.; and Ommer, B. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Zero-shot text-to-image generation
Ramesh, A.; Pavlov, M.; Goh, G.; Gray, S.; Voss, C.; Radford, A.; Chen, M.; and Sutskever, I. 2021 · 2021
Cited alongside, same era.
Layouttransformer: Scene layout generation with conceptual and spatial diversity
Yang, C.-F.; Fan, W.-C.; Yang, F.-E.; and Wang, Y.-C. F. 2021 · 2021
Cited alongside, same era.
ediffi: Text-to-image diffusion models with an ensemble of expert denoisers
Balaji, Y.; Nah, S.; Huang, X.; Vahdat, A.; Song, J.; Kreis, K.; Aittala, M.; Aila, T.; Laine, S.; Catanzaro, B.; et al. 2022 · 2022
Cited alongside, same era.
Shape-Guided Diffusion with Inside-Outside Attention
Huk Park, D.; Luo, G.; Toste, C.; Azadi, S.; Liu, X.; Karalashvili, M.; Rohrbach, A.; and Darrell, T. 2022 · 2022
Cited alongside, same era.
Diffusion models in vision: A survey
Croitoru, F.-A.; Hondru, V.; Ionescu, R. T.; and Shah, M. 2023 · 2023
Closest in time.
Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis
Feng, W.; He, X.; Fu, T.-J.; Jampani, V.; Akula, A. R.; Narayana, P.; Basu, S.; Wang, X. E.; and Wang, W. Y. 2023 · 2023
Closest in time.
SVDiff: Compact Parameter Space for Diffusion Fine-Tuning
Han, L.; Li, Y.; Zhang, H.; Milanfar, P.; Metaxas, D.; and Yang, F. 2023 · 2023
Closest in time.
Mixture of Diffusers for scene composition and high resolution image generation
Jiménez, Á. B. 2023 · 2023
Closest in time.
Gligen: Open-set grounded text-to-image generation
Li, Y.; Liu, H.; Wu, Q.; Mu, F.; Yang, J.; Gao, J.; Li, C.; and Lee, Y. J. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Compositional visual generation with composable diffusion models
Liu, N.; Li, S.; Du, Y.; Torralba, A.; and Tenenbaum, J. B. 2022 · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022 · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E. L.; Ghasemipour, K.; Gontijo Lopes, R.; Karagol Ayan, B.; Salimans, T.; et al. 2022 · 2022
Cited alongside, same era.
Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models
Chefer, H.; Alaluf, Y.; Vinker, Y.; Wolf, L.; and Cohen-Or, D. 2023 · 2023
Cited alongside, same era.
Training-Free Layout Control with Cross-Attention Guidance
Chen, M.; Laina, I.; and Vedaldi, A. 2023 · 2023
Cited alongside, same era.
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Liu, S.; Zeng, Z.; Ren, T.; Li, F.; Zhang, H.; Yang, J.; Li, C.; Yang, J.; Su, H.; Zhu, J.; et al. 2023a
Cited in the paper.
Lian, L.; Li, B.; Yala, A.; and Darrell, T. 2023 · 2023
Closest in time.
Directed Diffusion: Direct Control of Object Placement through Attention Guidance
Ma, W.-D. K.; Lewis, J.; Kleijn, W. B.; and Leung, T. 2023 · 2023
Closest in time.
Harnessing the spatial-temporal attention of diffusion models for high-fidelity text-to-image synthesis
Wu, Q.; Liu, Y.; Zhao, H.; Bui, T.; Lin, Z.; Zhang, Y.; and Chang, S. 2023 · 2023
Closest in time.
Diffusion models: A comprehensive survey of methods and applications
Yang, L.; Zhang, Z.; Song, Y.; Hong, S.; Xu, R.; Zhao, Y.; Zhang, W.; Cui, B.; and Yang, M.-H. 2023 · 2023
Closest in time.
Adding conditional control to text-to-image diffusion models
Zhang, L.; Rao, A.; and Agrawala, M. 2023 · 2023
Closest in time.