Fetching the paper…
Reading the bibliography…
Recent large-scale generative models learned on big data are capable of synthesizing incredible images yet suffer from limited controllability.
Aspects of the Theory of Syntax
Chomsky, N · 1965
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S · 2010
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J. N., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2011
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S. J., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M. S., Berg, A. C., and Fei-Fei, L · 2014
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson, J., Hariharan, B., van der Maaten, L., Fei-Fei, L., Zitnick, C. L., and Girshick, R. B · 2016
Earlier work this paper cites.
Rayleigh: Search image collections by multiple color palettes or by image color similarity., 2016
Sergeyk · 2016
Earlier work this paper cites.
Building machines that learn and think like people
Lake, B. M., Ullman, T. D., Tenenbaum, J. B., and Gershman, S. J · 2017
Earlier work this paper cites.
Webvision database: Visual learning and understanding from web data
Li, W., Wang, L., Li, W., Agustsson, E., and Gool, L. V · 2017
Earlier work this paper cites.
Mastering sketching: Adversarial augmentation for structured prediction
Simo-Serra, E., Iizuka, S., and Ishikawa, H · 2017
Earlier work this paper cites.
AttnGAN: Fine-grained text to image generation with attentional generative adversarial networks
Xu, T., Zhang, P., Huang, Q., Zhang, H., Gan, Z., Huang, X., and He, X · 2017
Earlier work this paper cites.
Measuring compositional generalization: A comprehensive method on realistic data
Keysers, D., Schärli, N., Scales, N., Buisman, H., Furrer, D., Kashubin, S., Momchev, N., Sinopalnikov, D., Stafiniak, L., Tihon, T., Tsarkov, D., Wang, X., van Zee, M., and Bousquet, O · 2019
Earlier work this paper cites.
Image synthesis from reconfigurable layout and style
Sun, W. and Wu, T · 2019
Earlier work this paper cites.
DM-GAN: Dynamic memory generative adversarial networks for text-to-image synthesis
Zhu, M., Pan, P., Chen, W., and Yang, Y · 2019
Earlier work this paper cites.
Taming Transformers for high-resolution image synthesis
Esser, P., Rombach, R., and Ommer, B · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
YOLOv5 by Ultralytics, 2020
Jocher, G · 2020
Cited alongside, same era.
Improved techniques for training score-based generative models
Song, Y. and Ermon, S · 2020
Cited alongside, same era.
Diffusion models beat GANs on image synthesis
Dhariwal, P. and Nichol, A · 2021
Cited alongside, same era.
CogView: Mastering text-to-image generation via Transformers
Ding, M., Yang, Z., Hong, W., Zheng, W., Zhou, C., Yin, D., Lin, J., Zou, X., Shao, Z., Yang, H., and Tang, J · 2021
Cited alongside, same era.
Multimodal conditional image synthesis with product-of-experts GANs
Huang, X., Mallya, A., Wang, T.-C., and Liu, M.-Y · 2021
Cited alongside, same era.
Compositional visual generation with composable diffusion models
Liu, N., Li, S., Du, Y., Torralba, A., and Tenenbaum, J. B · 2022
Later among the works it cites.
Null-text inversion for editing real images using guided diffusion models
Mokady, R., Hertz, A., Aberman, K., Pritch, Y., and Cohen-Or, D · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Later among the works it cites.
Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer
Ranftl, R., Lasinger, K., Hafner, D., Schindler, K., and Koltun, V · 2022
Later among the works it cites.
DreamBooth: Fine tuning text-to-image diffusion models for subject-driven generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nichol, A. and Dhariwal, P · 2021
Cited alongside, same era.
GLIDE: Towards photorealistic image generation and editing with text-guided diffusion models
Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., and Chen, M · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2021
Cited alongside, same era.
Pixel difference networks for efficient edge detection
Su, Z., Liu, W., Yu, Z., Hu, D., Liao, Q., Tian, Q., Pietikäinen, M., and Liu, L · 2021
Cited alongside, same era.
Vector-quantized image modeling with improved VQGAN
Yu, J., Li, X., Koh, J. Y., Zhang, H., Pang, R., Qin, J., Ku, A., Xu, Y., Baldridge, J., and Wu, Y · 2021
Cited alongside, same era.
Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., and Aberman, K · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., Salimans, T., Ho, J., Fleet, D. J., and Norouzi, M · 2022
Later among the works it cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., Schramowski, P., Kundurthy, S., Crowson, K., Schmidt, L., Kaczmarczyk, R., and Jitsev, J · 2022
Later among the works it cites.
Stable diffusion 2.0 release., 2022
stability.ai · 2022
Later among the works it cites.
Sketch-guided text-to-image diffusion models
Voynov, A., Aberman, K., and Cohen-Or, D · 2022
Later among the works it cites.
EDICT: Exact diffusion inversion via coupled transformations
Wallace, B., Gokul, A., and Naik, N · 2022
Later among the works it cites.
SmartBrush: Text and shape guided object inpainting with diffusion model
Xie, S., Zhang, Z., Lin, Z., Hinz, T., and Zhang, K · 2022
Later among the works it cites.
Diffusion-based scene graph to image generation with masked contrastive pre-training
Yang, L., Huang, Z., Song, Y., Hong, S., Li, G., Zhang, W., Cui, B., Ghanem, B., and Yang, M.-H · 2022
Later among the works it cites.
Scaling autoregressive models for content-rich text-to-image generation
Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., Hutchinson, B. C., Han, W., Parekh, Z., Li, X., Zhang, H., Baldridge, J., and Wu, Y · 2022
Later among the works it cites.
Muse: Text-to-image generation via masked generative transformers
Chang, H., Zhang, H., Barber, J., Maschinot, A., Lezama, J., Jiang, L., Yang, M., Murphy, K. P., Freeman, W. T., Rubinstein, M., Li, Y., and Krishnan, D · 2023
Closest in time.
GLIGEN: Open-set grounded text-to-image generation
Li, Y., Liu, H., Wu, Q., Mu, F., Yang, J., Gao, J., Li, C., and Lee, Y. J · 2023
Closest in time.