Fetching the paper…
Reading the bibliography…
In recent years, text-to-image (T2I) generation models have made significant progress in generating high-quality images that align with text descriptions.
Auto-encoding variational bayes
Kingma, D. P · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Multimodal datasets: misogyny, pornography, and malignant stereotypes
Birhane, A., Prabhu, V. U., and Kahembwe, E · 2021
Earlier work this paper cites.
Machine unlearning
Bourtoule, L., Chandrasekaran, V., Choquette-Choo, C. A., Jia, H., Travers, A., Zhang, B., Lie, D., and Papernot, N · 2021
Earlier work this paper cites.
Datasheets for datasets
Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Iii, H. D., and Crawford, K · 2021
Earlier work this paper cites.
Knowledge distillation: A survey
Gou, J., Yu, B., Maybank, S. J., and Tao, D · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S · 2021
Earlier work this paper cites.
Red-teaming the stable diffusion safety filter
Rando, J., Paleka, D., Lindner, D., Heim, L., and Tramer, F · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Earlier work this paper cites.
Can machines help us answering question 16 in datasheets, and in turn reflecting on inappropriate content?
Schramowski, P., Tauchmann, C., and Kersting, K · 2022
Earlier work this paper cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al · 2022
Earlier work this paper cites.
Make-a-video: Text-to-video generation without text-video data
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., et al · 2022
Earlier work this paper cites.
Improving image generation with better captions
Betker, J., Goh, G., Jing, L., Brooks, T., Wang, J., Li, L., Ouyang, L., Zhuang, J., Lee, J., Guo, Y., et al · 2023
Cited alongside, same era.
Easily accessible text-to-image generation amplifies demographic stereotypes at large scale
Bianchi, F., Kalluri, P., Durmus, E., Ladhak, F., Cheng, M., Nozza, D., Hashimoto, T., Jurafsky, D., Zou, J., and Caliskan, A · 2023
Cited alongside, same era.
Erasing concepts from diffusion models
Gandikota, R., Materzynska, J., Fiotto-Kaufman, J., and Bau, D · 2023
Cited alongside, same era.
Ablating concepts in text-to-image diffusion models
Kumari, N., Zhang, B., Wang, S.-Y., Shechtman, E., Zhang, R., and Zhu, J.-Y · 2023
Cited alongside, same era.
Circumventing concept erasure methods for text-to-image generative models
Pham, M., Marshall, K. O., Cohen, N., Mittal, G., and Hegde, C · 2023
Cited alongside, same era.
Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models
One-dimensional adapter to rule them all: Concepts diffusion models and erasing applications
Lyu, M., Yang, Y., Hong, H., Chen, H., Jin, X., He, Y., Xue, H., Han, J., and Ding, G · 2024
Closest in time.
Safety checker model card
Machine Vision & Learning Group LMU · 2024
Closest in time.
Safe-clip: Removing nsfw concepts from vision-and-language models
Poppi, S., Poppi, T., Cocchi, F., Cornia, M., Baraldi, L., Cucchiara, R., et al · 2024
Closest in time.
Cola: A benchmark for compositional text-to-image retrieval
Ray, A., Radenovic, F., Dubey, A., Plummer, B., Krishna, R., and Saenko, K · 2024
Closest in time.
Stable diffusion v1 model card
stability ai · 2024
Closest in time.
Stable diffusion v2
stability ai · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Qu, Y., Shen, X., He, X., Backes, M., Zannettou, S., and Zhang, Y · 2023
Cited alongside, same era.
Safe latent diffusion: Mitigating inappropriate degeneration in diffusion models
Schramowski, P., Brack, M., Deiseroth, B., and Kersting, K · 2023
Cited alongside, same era.
Rickrolling the artist: Injecting backdoors into text encoders for text-to-image synthesis
Struppek, L., Hintersdorf, D., and Kersting, K · 2023
Cited alongside, same era.
Sneakyprompt: Jailbreaking text-to-image generative models
Yang, Y., Hui, B., Yuan, H., Gong, N., and Cao, Y · 2023
Cited alongside, same era.
People are creating an average of 34 million images per day. statistics for 2024
EveryPixel · 2024
Cited alongside, same era.
Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., et al · 2024
Cited alongside, same era.
A detailed review on word embedding techniques with emphasis on word2vec
Johnson, S. J., Murty, M. R., and Navakanth, I · 2024
Cited alongside, same era.
stability ai · 2024
Closest in time.
Ring-a-bell! how reliable are concept removal methods for diffusion models?
Tsai, Y.-L., Hsu, C.-Y., Xie, C., Lin, C.-H., Chen, J. Y., Li, B., Chen, P.-Y., Yu, C.-M., and Huang, C.-Y · 2024
Closest in time.
Moderator: Moderating text-to-image diffusion models through fine-grained context-based policies
Wang, P., Li, Q., Yu, L., Wang, Z., Li, A., and Jin, H · 2024
Closest in time.
Universal prompt optimizer for safe text-to-image generation
Wu, Z., Gao, H., Wang, Y., Zhang, X., and Wang, S · 2024
Closest in time.
Mma-diffusion: Multimodal attack on diffusion models
Yang, Y., Gao, R., Wang, X., Ho, T.-Y., Xu, N., and Xu, Q · 2024
Closest in time.
Guardt2i: Defending text-to-image models from adversarial prompts
Yang, Y., Gao, R., Yang, X., Zhong, J., and Xu, Q · 2024
Closest in time.
Evaluating semantic variation in text-to-image synthesis: A causal perspective
Zhu, X., Sun, P., Song, Y., Xiao, Y., Li, Z., Wang, C., Huang, J., Yang, B., and Xu, X · 2024
Closest in time.
Diffit: Diffusion vision transformers for image generation
Hatamizadeh, A., Song, J., Liu, G., Kautz, J., and Vahdat, A · 2025
Closest in time.
Latent guard: a safety framework for text-to-image generation
Liu, R., Khakzar, A., Gu, J., Chen, Q., Torr, P., and Pizzati, F · 2025
Closest in time.