Fetching the paper…
Reading the bibliography…
Recent text-to-image (T2I) models have had great success, and many benchmarks have been proposed to evaluate their performance and safety.
Siren’s song in the ai ocean: A survey on hallucination in large language models
Zhang, Y., Li, Y., Cui, L., Cai, D., Liu, L., Fu, T., Huang, X., Zhao, E., Zhang, Y., Chen, Y., et al · 2000
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
Wah, C., Branson, S., Welinder, P., Perona, P., and Belongie, S · 2011
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Deep learning face representation by joint identification-verification
Sun, Y., Chen, Y., Wang, X., and Tang, X · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Improved techniques for training gans
Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., and Chen, X · 2016
Earlier work this paper cites.
Joint face detection and alignment using multitask cascaded convolutional networks
Zhang, K., Zhang, Z., Li, Z., and Qiao, Y · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Generative adversarial networks: An overview
Creswell, A., White, T., Dumoulin, V., Arulkumaran, K., Sengupta, B., and Bharath, A. A · 2018
Earlier work this paper cites.
Attngan: Fine-grained text to image generation with attentional generative adversarial networks
Xu, T., Zhang, P., Huang, Q., Zhang, H., Gan, Z., Huang, X., and He, X · 2018
Earlier work this paper cites.
Arcface: Additive angular margin loss for deep face recognition
Deng, J., Guo, J., Xue, N., and Zafeiriou, S · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Clipscore: A reference-free evaluation metric for image captioning
Hessel, J., Holtzman, A., Forbes, M., Bras, R. L., and Choi, Y · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Earlier work this paper cites.
GitHub Merge: [Safety Checker] Add Safety Checker Module
CompVis · 2022
Earlier work this paper cites.
Cogview2: Faster and better text-to-image generation via hierarchical transformers
Ding, M., Zheng, W., Hong, W., and Tang, J · 2022
Cited alongside, same era.
Make-a-scene: Scene-based text-to-image generation with human priors
Gafni, O., Polyak, A., Ashual, O., Sheynin, S., Parikh, D., and Taigman, Y · 2022
Cited alongside, same era.
An image is worth one word: Personalizing text-to-image generation using textual inversion
Gal, R., Alaluf, Y., Atzmon, Y., Patashnik, O., Bermano, A. H., Chechik, G., and Cohen-Or, D · 2022
Cited alongside, same era.
Human evaluation of text-to-image models on a multi-task benchmark
Petsiuk, V., Siemenn, A. E., Surbehera, S., Chin, Z., Tyser, K., Hunter, G., Raghavan, A., Hicke, Y., Plummer, B. A., Kerret, O., et al · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Stable bias: Analyzing societal representations in diffusion models
Luccioni, A. S., Akiki, C., Mitchell, M., and Jernite, Y · 2023
Later among the works it cites.
Gpt-4v(ision) system card
OpenAI · 2023
Later among the works it cites.
Dall·e 3 system card, 2023
OpenAI · 2023
Later among the works it cites.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R · 2023
Later among the works it cites.
Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models
Qu, Y., Shen, X., He, X., Backes, M., Zannettou, S., and Zhang, Y · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Cited alongside, same era.
Red-teaming the stable diffusion safety filter
Rando, J., Paleka, D., Lindner, D., Heim, L., and Tramèr, F · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al · 2022
Cited alongside, same era.
Df-gan: A simple and effective baseline for text-to-image synthesis
Tao, M., Tang, H., Wu, F., Jing, X.-Y., Bao, B.-K., and Xu, C · 2022
Cited alongside, same era.
Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models
Bakr, E. M., Sun, P., Shen, X., Khan, F. F., Li, L. E., and Elhoseiny, M · 2023
Cited alongside, same era.
Multifusion: Fusing pre-trained models for multi-lingual, multi-modal image generation
Bellagente, M., Brack, M., Teufel, H., Friedrich, F., Deiseroth, B., Eichenberg, C., Dai, A., Baldock, R., Nanda, S., Oostermeijer, K., et al · 2023
Cited alongside, same era.
Improving image generation with better captions
Betker, J., Goh, G., Jing, L., Brooks, T., Wang, J., Li, L., Ouyang, L., Zhuang, J., Lee, J., Guo, Y., et al · 2023
Cited alongside, same era.
Safe latent diffusion: Mitigating inappropriate degeneration in diffusion models
Schramowski, P., Brack, M., Deiseroth, B., and Kersting, K · 2023
Later among the works it cites.
The bias amplification paradox in text-to-image generation
Seshadri, P., Singh, S., and Elazar, Y · 2023
Later among the works it cites.
Imagereward: Learning and evaluating human preferences for text-to-image generation
Xu, J., Liu, X., Wu, Y., Tong, Y., Li, Q., Ding, M., Tang, J., and Dong, Y · 2023
Later among the works it cites.
Civitai: The home of open-source generative ai, 2024
Civitai · 2024
Closest in time.
Bard, 2024
Google · 2024
Closest in time.
Optimizing prompts for text-to-image generation
Hao, Y., Chi, Z., Dong, L., and Wei, F · 2024
Closest in time.
Lexica, 2024
Lexica · 2024
Closest in time.
Safety checker model card
Machine Vision & Learning Group LMU · 2024
Closest in time.
Midjourney, 2023
Midjourney, I · 2024
Closest in time.
Openai content pilicy, 2024
OpenAI · 2024
Closest in time.
Eyes wide shut? exploring the visual shortcomings of multimodal llms
Tong, S., Liu, Z., Zhai, Y., Ma, Y., LeCun, Y., and Xie, S · 2024
Closest in time.