Fetching the paper…
Reading the bibliography…
Despite significant progress in Text-to-Image (T2I) generative models, even lengthy and complex text descriptions still struggle to convey detailed controls.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Improved techniques for training gans
Salimans, T.; Goodfellow, I.; Zaremba, W.; Cheung, V.; Radford, A.; and Chen, X. 2016 · 2016
Earlier work this paper cites.
Generative visual manipulation on the natural image manifold
Zhu, J.-Y.; Krähenbühl, P.; Shechtman, E.; and Efros, A. A. 2016 · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; and Hochreiter, S. 2017 · 2017
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks
Isola, P.; Zhu, J.-Y.; Zhou, T.; and Efros, A. A. 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A.; Vinyals, O.; et al. 2017 · 2017
Earlier work this paper cites.
Coco-stuff: Thing and stuff classes in context
Caesar, H.; Uijlings, J.; and Ferrari, V. 2018 · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R.; Isola, P.; Efros, A. A.; Shechtman, E.; and Wang, O. 2018 · 2018
Earlier work this paper cites.
Semantic image synthesis with spatially-adaptive normalization
Park, T.; Liu, M.-Y.; Wang, T.-C.; and Zhu, J.-Y. 2019 · 2019
Earlier work this paper cites.
Image synthesis from reconfigurable layout and style
Sun, W.; and Wu, T. 2019 · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Earlier work this paper cites.
Maskgan: Towards diverse and interactive facial image manipulation
Lee, C.-H.; Liu, Z.; Wu, L.; and Luo, P. 2020 · 2020
Earlier work this paper cites.
Controllable Person Image Synthesis with Attribute-Decomposed GAN
Men, Y.; Mao, Y.; Jiang, Y.; Ma, W.-Y.; and Lian, Z. 2020 · 2020
Earlier work this paper cites.
Denoising Diffusion Implicit Models
Song, J.; Meng, C.; and Ermon, S. 2020 · 2020
Cited alongside, same era.
Diffusion models beat gans on image synthesis
Dhariwal, P.; and Nichol, A. 2021 · 2021
Cited alongside, same era.
Taming transformers for high-resolution image synthesis
Esser, P.; Rombach, R.; and Ommer, B. 2021 · 2021
Cited alongside, same era.
Context-aware layout to image generation with enhanced object appearance
He, S.; Liao, W.; Yang, M. Y.; Yang, Y.; Song, Y.-Z.; Rosenhahn, B.; and Xiang, T. 2021 · 2021
Cited alongside, same era.
Image Synthesis from Layout with Locality-Aware Mask Adaption
Li, Z.; Wu, J.; Koh, I.; Tang, Y.; and Sun, L. 2021 · 2021
Cited alongside, same era.
Nerf: Representing scenes as neural radiance fields for view synthesis
Mildenhall, B.; Srinivasan, P. P.; Tancik, M.; Barron, J. T.; Ramamoorthi, R.; and Ng, R. 2021 · 2021
Cited alongside, same era.
Stable Diffusion with Diffusers
Patil, S.; Cuenca, P.; Lambert, N.; and von Platen, P. 2022 · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022 · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E. L.; Ghasemipour, K.; Gontijo Lopes, R.; Karagol Ayan, B.; Salimans, T.; et al. 2022 · 2022
Later among the works it cites.
Pretraining is all you need for image-to-image translation
Wang, T.; Zhang, T.; Zhang, B.; Ouyang, H.; Chen, D.; Chen, Q.; and Wen, F. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Encoding in style: a stylegan encoder for image-to-image translation
Richardson, E.; Alaluf, Y.; Patashnik, O.; Nitzan, Y.; Azar, Y.; Shapiro, S.; and Cohen-Or, D. 2021 · 2021
Cited alongside, same era.
Score-Based Generative Modeling through Stochastic Differential Equations
Song, Y.; Sohl-Dickstein, J.; Kingma, D. P.; Kumar, A.; Ermon, S.; and Poole, B. 2021 · 2021
Cited alongside, same era.
Learning layout and style reconfigurable gans for controllable image synthesis
Sun, W.; and Wu, T. 2021 · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; et al. 2022 · 2022
Cited alongside, same era.
Classifier-free diffusion guidance
Ho, J.; and Salimans, T. 2022 · 2022
Cited alongside, same era.
Modeling image composition for complex scene generation
Yang, Z.; Liu, D.; Wang, C.; Yang, J.; and Tao, D. 2022 · 2022
Later among the works it cites.
Diffusion Models in Vision: A Survey
Croitoru, F.-A.; Hondru, V.; Ionescu, R. T.; and Shah, M. 2023 · 2023
Closest in time.
Gligen: Open-set grounded text-to-image generation
Li, Y.; Liu, H.; Wu, Q.; Mu, F.; Yang, J.; Gao, J.; Li, C.; and Lee, Y. J. 2023 · 2023
Closest in time.
Freestyle Layout-to-Image Synthesis
Xue, H.; Huang, Z.; Sun, Q.; Song, L.; and Zhang, W. 2023 · 2023
Closest in time.
Reco: Region-controlled text-to-image generation
Yang, Z.; Wang, J.; Gan, Z.; Li, L.; Lin, K.; Wu, C.; Duan, N.; Liu, Z.; Liu, C.; Zeng, M.; et al. 2023 · 2023
Closest in time.
Adding conditional control to text-to-image diffusion models
Zhang, L.; and Agrawala, M. 2023 · 2023
Closest in time.
LayoutDiffusion: Controllable Diffusion Model for Layout-to-image Generation
Zheng, G.; Zhou, X.; Li, X.; Qi, Z.; Shan, Y.; and Li, X. 2023 · 2023
Closest in time.
TryOnDiffusion: A Tale of Two UNets
Zhu, L.; Yang, D.; Zhu, T.; Reda, F.; Chan, W.; Saharia, C.; Norouzi, M.; and Kemelmacher-Shlizerman, I. 2023 · 2023
Closest in time.