Fetching the paper…
Reading the bibliography…
In the evolving domain of text-to-image generation, diffusion models have emerged as powerful tools in content creation.
Arbitrary style transfer in real-time with adaptive instance normalization
Huang, X. and Belongie, S · 2017
Earlier work this paper cites.
Nsml: A machine learning platform that enables you to focus on your models
Sung, N., Kim, M., Jo, H., Yang, Y., Kim, J., Lausen, L., Kim, Y., Lee, G., Kwak, D., Ha, J.-W., et al · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Nsml: Meet the mlaas platform with a real-world case study
Kim, H., Kim, M., Seo, D., Kim, J., Park, H., Park, S., Jo, H., Kim, K., Yang, Y., Kim, Y., et al · 2018
Earlier work this paper cites.
Avatar-net: Multi-scale zero-shot style transfer by feature decoration
Sheng, L., Lin, Z., Shao, J., and Wang, X · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O · 2018
Earlier work this paper cites.
Arbitrary style transfer with style-attentional networks
Park, D. Y. and Lee, K. H · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S · 2020
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A · 2021
Earlier work this paper cites.
Adaattn: Revisit attention mechanism in arbitrary neural style transfer
Liu, S., Lin, T., He, D., Li, F., Wang, M., Li, X., Sun, Z., Li, Q., and Ding, E · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2021
Earlier work this paper cites.
An image is worth one word: Personalizing text-to-image generation using textual inversion
Gal, R., Alaluf, Y., Atzmon, Y., Patashnik, O., Bermano, A. H., Chechik, G., and Cohen-or, D · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al · 2022
Cited alongside, same era.
Break-a-scene: Extracting multiple concepts from a single image
Avrahami, O., Aberman, K., Fried, O., Cohen-Or, D., and Lischinski, D · 2023
Cited alongside, same era.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R · 2023
Later among the works it cites.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., and Aberman, K · 2023
Later among the works it cites.
Low-rank adaptation for fast text-to-image diffusion fine-tuning
Ryu, S · 2023
Later among the works it cites.
Styledrop: Text-to-image synthesis of any style
Sohn, K., Jiang, L., Barber, J., Lee, K., Ruiz, N., Krishnan, D., Chang, H., Li, Y., Essa, I., Rubinstein, M., Hao, Y., Entis, G., Blok, I., and Chin, D. C · 2023
Later among the works it cites.
Diffusion image analogies
Šubrtová, A., Lukáč, M., Čech, J., Futschik, D., Shechtman, E., and Sỳkora, D · 2023
Later among the works it cites.
Imagebrush: Learning visual in-context instructions for exemplar-based image manipulation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing
Cao, M., Wang, X., Qi, Z., Shan, Y., Qie, X., and Zheng, Y · 2023
Cited alongside, same era.
Diffusion in style
Everaert, M. N., Bocchio, M., Arpa, S., Süsstrunk, S., and Achanta, R · 2023
Cited alongside, same era.
Highly personalized text embedding for image manipulation by stable diffusion
Han, I., Yang, S., Kwon, T., and Ye, J. C · 2023
Cited alongside, same era.
Style aligned image generation via shared attention
Hertz, A., Voynov, A., Fruchter, S., and Cohen-Or, D · 2023
Cited alongside, same era.
Multi-concept customization of text-to-image diffusion
Kumari, N., Zhang, B., Zhang, R., Shechtman, E., and Zhu, J.-Y · 2023
Cited alongside, same era.
Diffusion models already have a semantic latent space
Kwon, M., Jeong, J., and Uh, Y · 2023
Cited alongside, same era.
Gligen: Open-set grounded text-to-image generation
Li, Y., Liu, H., Wu, Q., Mu, F., Yang, J., Gao, J., Li, C., and Lee, Y. J · 2023
Cited alongside, same era.
Sun, Y., Yang, Y., Peng, H., Shen, Y., Yang, Y., Hu, H., Qiu, L., and Koike, H · 2023
Later among the works it cites.
Plug-and-play diffusion features for text-driven image-to-image translation
Tumanyan, N., Geyer, M., Bagon, S., and Dekel, T · 2023
Later among the works it cites.
p + p+ : Extended textual conditioning in text-to-image generation
Voynov, A., Chu, Q., Cohen-Or, D., and Aberman, K · 2023
Later among the works it cites.
Styleadapter: A single-pass lora-free model for stylized image generation
Wang, Z., Wang, X., Xie, L., Qi, Z., Shan, Y., Wang, W., and Luo, P · 2023
Later among the works it cites.
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
Wu, J. Z., Ge, Y., Wang, X., Lei, S. W., Gu, Y., Shi, Y., Hsu, W., Shan, Y., Qie, X., and Shou, M. Z · 2023
Later among the works it cites.
Rerender a video: Zero-shot text-guided video-to-video translation
Yang, S., Zhou, Y., Liu, Z., and Loy, C. C · 2023
Later among the works it cites.
Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models
Ye, H., Zhang, J., Liu, S., Han, X., and Yang, W · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
Zhang, L., Rao, A., and Agrawala, M · 2023
Later among the works it cites.