Fetching the paper…
Reading the bibliography…
Despite the tremendous success in text-to-image generative models, localized text-to-image generation (that is, generating objects or features at specific locations in an image while maintaining a consistent overall generation) still requires either explicit training or substantial additional inference time.
Im2text: Describing images using 1 million captioned photographs
V. Ordonez, G. Kulkarni, and T. Berg · 2011
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Conditional generative adversarial nets, 2014
M. Mirza and S. Osindero · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context, 2015
T.-Y. Lin, M. Maire, S. Belongie, L. Bourdev, R. Girshick, J. Hays, P. Perona, D. Ramanan, C. L. Zitnick, and P. Dollár · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics, 2015
J. Sohl-Dickstein, E. A. Weiss, N. Maheswaranathan, and S. Ganguli · 2015
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations, 2016
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, M. S. Bernstein, and F.-F. Li · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks
P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros · 2017
Earlier work this paper cites.
Demystifying MMD GANs
M. Bińkowski, D. J. Sutherland, M. Arbel, and A. Gretton · 2018
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks, 2018
T. Karras, S. Laine, and T. Aila · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
P. Sharma, N. Ding, S. Goodman, and R. Soricut · 2018
Earlier work this paper cites.
Semantic image synthesis with spatially-adaptive normalization
T. Park, M.-Y. Liu, T.-C. Wang, and J.-Y. Zhu · 2019
Cited alongside, same era.
Objects365: A large-scale, high-quality dataset for object detection
S. Shao, Z. Li, T. Zhang, C. Peng, G. Yu, X. Zhang, J. Li, and J. Sun · 2019
Cited alongside, same era.
Image synthesis from reconfigurable layout and style, 2019
W. Sun and T. Wu · 2019
Cited alongside, same era.
Image generation from layout, 2019
B. Zhao, L. Meng, W. Yin, and L. Sigal · 2019
Cited alongside, same era.
Taming transformers for high-resolution image synthesis, 2020
P. Esser, R. Rombach, and B. Ommer · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Auto-encoding variational bayes, 2022
D. P. Kingma and M. Welling · 2022
Later among the works it cites.
Grounded language-image pre-training
L. H. Li*, P. Zhang*, H. Zhang*, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J.-N. Hwang, K.-W. Chang, and J. Gao · 2022
Later among the works it cites.
Compositional visual generation with composable diffusion models
N. Liu, S. Li, Y. Du, A. Torralba, and J. B. Tenenbaum · 2022
Later among the works it cites.
SDEdit: Guided image synthesis and editing with stochastic differential equations
C. Meng, Y. He, Y. Song, J. Song, J. Wu, J.-Y. Zhu, and S. Ermon · 2022
Later among the works it cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models, 2022
A. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding, 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Diffusion models beat gans on image synthesis
P. Dhariwal and A. Nichol · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision, 2021
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Cited alongside, same era.
Zero-shot text-to-image generation, 2021
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models, 2021
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2021
Cited alongside, same era.
Designing an encoder for stylegan image manipulation
O. Tov, Y. Alaluf, Y. Nitzan, O. Patashnik, and D. Cohen-Or · 2021
Cited alongside, same era.
Training-free structured diffusion guidance for compositional text-to-image synthesis
W. Feng, X. He, T.-J. Fu, V. Jampani, A. Akula, P. Narayana, S. Basu, X. E. Wang, and W. Y. Wang · 2022
Cited alongside, same era.
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. Denton, S. K. S. Ghasemipour, B. K. Ayan, S. S. Mahdavi, R. G. Lopes, T. Salimans, J. Ho, D. J. Fleet, and M. Norouzi · 2022
Later among the works it cites.
Winoground: Probing vision and language models for visio-linguistic compositionality
T. Thrush, R. Jiang, M. Bartolo, A. Singh, A. Williams, D. Kiela, and C. Ross · 2022
Later among the works it cites.
Multidiffusion: Fusing diffusion paths for controlled image generation
O. Bar-Tal, L. Yariv, Y. Lipman, and T. Dekel · 2023
Closest in time.
Semantic segment anything
J. Chen, Z. Yang, and L. Zhang · 2023
Closest in time.
YOLO by Ultralytics, Jan. 2023
G. Jocher, A. Chaurasia, and J. Qiu · 2023
Closest in time.
Scaling up gans for text-to-image synthesis
M. Kang, J.-Y. Zhu, R. Zhang, J. Park, E. Shechtman, S. Paris, and T. Park · 2023
Closest in time.
Segment anything, 2023
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. Girshick · 2023
Closest in time.
Gligen: Open-set grounded text-to-image generation
Y. Li, H. Liu, Q. Wu, F. Mu, J. Yang, J. Gao, C. Li, and Y. J. Lee · 2023
Closest in time.