Fetching the paper…
Reading the bibliography…
Existing text-to-image models struggle to follow complex text prompts, raising the need for extra grounding inputs for better controllability.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Deep learning in object recognition, detection, and segmentation
Wang, X. et al · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2018
Earlier work this paper cites.
Compositional visual generation with energy based models
Du, Y., Li, S., and Mordatch, I · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2020
Earlier work this paper cites.
Fourier features let networks learn high frequency functions in low dimensional domains
Tancik, M., Srinivasan, P., Mildenhall, B., Fridovich-Keil, S., Raghavan, N., Singhal, U., Ramamoorthi, R., Barron, J., and Ng, R · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Dhariwal, P. and Nichol, A · 2021
Earlier work this paper cites.
Cogview: Mastering text-to-image generation via transformers
Ding, M., Yang, Z., Hong, W., Zheng, W., Zhou, C., Yin, D., Lin, J., Zou, X., Shao, Z., Yang, H., et al · 2021
Earlier work this paper cites.
Controllable and compositional generation with latent-space energy-based models
Nie, W., Vahdat, A., and Anandkumar, A · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Earlier work this paper cites.
Blended diffusion for text-driven editing of natural images
Avrahami, O., Lischinski, D., and Fried, O · 2022
Earlier work this paper cites.
ediffi: Text-to-image diffusion models with an ensemble of expert denoisers
Balaji, Y., Nah, S., Huang, X., Vahdat, A., Song, J., Kreis, K., Aittala, M., Aila, T., Laine, S., Catanzaro, B., et al · 2022
Earlier work this paper cites.
Re-imagen: Retrieval-augmented text-to-image generator
Chen, W., Hu, H., Saharia, C., and Cohen, W. W · 2022
Earlier work this paper cites.
Blobgan: Spatially disentangled scene representations
Epstein, D., Park, T., Zhang, R., Shechtman, E., and Efros, A. A · 2022
Cited alongside, same era.
Training-free structured diffusion guidance for compositional text-to-image synthesis
Feng, W., He, X., Fu, T.-J., Jampani, V., Akula, A. R., Narayana, P., Basu, S., Wang, X. E., and Wang, W. Y · 2022
Cited alongside, same era.
Make-a-scene: Scene-based text-to-image generation with human priors
Gafni, O., Polyak, A., Ashual, O., Sheynin, S., Parikh, D., and Taigman, Y · 2022
Cited alongside, same era.
Grounded language-image pre-training
Li, L. H., Zhang, P., Zhang, H., Yang, J., Li, C., Zhong, Y., Wang, L., Yuan, L., Zhang, L., Hwang, J.-N., et al · 2022
Cited alongside, same era.
Compositional visual generation with composable diffusion models
Liu, N., Li, S., Du, Y., Torralba, A., and Tenenbaum, J. B · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Diffusion self-guidance for controllable image generation
Epstein, D., Jabri, A., Poole, B., Efros, A. A., and Holynski, A · 2023
Later among the works it cites.
Expressive text-to-image generation with rich text
Ge, S., Park, T., Zhu, J.-Y., and Huang, J.-B · 2023
Later among the works it cites.
Generating images with multimodal language models
Koh, J. Y., Fried, D., and Salakhutdinov, R · 2023
Later among the works it cites.
Multi-concept customization of text-to-image diffusion
Kumari, N., Zhang, B., Zhang, R., Shechtman, E., and Zhu, J.-Y · 2023
Later among the works it cites.
Gligen: Open-set grounded text-to-image generation
Li, Y., Liu, H., Wu, Q., Mu, F., Yang, J., Gao, J., Li, C., and Lee, Y. J · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al · 2022
Cited alongside, same era.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al · 2022
Cited alongside, same era.
Knn-diffusion: Image generation via large-scale retrieval
Sheynin, S., Ashual, O., Polyak, A., Singer, U., Gafni, O., Nachmani, E., and Taigman, Y · 2022
Cited alongside, same era.
Scaling autoregressive models for content-rich text-to-image generation
Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., et al · 2022
Cited alongside, same era.
Lian, L., Li, B., Yala, A., and Darrell, T · 2023
Later among the works it cites.
Null-text inversion for editing real images using guided diffusion models
Mokady, R., Hertz, A., Aberman, K., Pritch, Y., and Cohen-Or, D · 2023
Later among the works it cites.
Grounded text-to-image synthesis with attention refocusing
Phung, Q., Ge, S., and Huang, J.-B · 2023
Later among the works it cites.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R · 2023
Later among the works it cites.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., and Aberman, K · 2023
Later among the works it cites.
Generative pretraining in multimodality
Sun, Q., Yu, Q., Cui, Y., Zhang, F., Zhang, X., Wang, Y., Gao, H., Liu, J., Huang, T., and Wang, X · 2023
Later among the works it cites.
Fastcomposer: Tuning-free multi-subject image generation with localized attention
Xiao, G., Yin, T., Freeman, W. T., Durand, F., and Han, S · 2023
Later among the works it cites.
Open-vocabulary panoptic segmentation with text-to-image diffusion models
Xu, J., Liu, S., Vahdat, A., Byeon, W., Wang, X., and De Mello, S · 2023
Later among the works it cites.
Reco: Region-controlled text-to-image generation
Yang, Z., Wang, J., Gan, Z., Li, L., Lin, K., Wu, C., Duan, N., Liu, Z., Liu, C., Zeng, M., et al · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
Zhang, L., Rao, A., and Agrawala, M · 2023
Later among the works it cites.
Layoutdiffusion: Controllable diffusion model for layout-to-image generation
Zheng, G., Zhou, X., Li, X., Qi, Z., Shan, Y., and Li, X · 2023
Later among the works it cites.