Fetching the paper…
Reading the bibliography…
This work focuses on generating high-quality images with specific style of reference images and content of provided textual descriptions.
IEEE Transactions on Visualization and Computer Graphics 10
Wang, B., Wang, W., Yang, H., Sun, J.: Efficient example-based painting and synthesis of 2d directional texture · 2004
Earlier work this paper cites.
IEEE Transactions on multimedia 15
Zhang, W., Cao, C., Chen, S., Liu, J., Tang, X.: Style transfer via image component analysis · 2013
Earlier work this paper cites.
arXiv preprint arXiv:1412.6980 (2014)
Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization · 2014
Earlier work this paper cites.
In: MICCAI (2015)
Ronneberger, O., Fischer, P., Brox, T.: U-net: Convolutional networks for biomedical image segmentation · 2015
Earlier work this paper cites.
In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2414–2423 (2016)
Gatys, L.A., Ecker, A.S., Bethge, M.: Image style transfer using convolutional neural networks · 2016
Earlier work this paper cites.
In: International conference on machine learning, pp. 1060–1069. PMLR (2016)
Reed, S., Akata, Z., Yan, X., Logeswaran, L., Schiele, B., Lee, H.: Generative adversarial text to image synthesis · 2016
Earlier work this paper cites.
In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3985–3993 (2017)
Gatys, L.A., Ecker, A.S., Bethge, M., Hertzmann, A., Shechtman, E.: Controlling perceptual factors in neural style transfer · 2017
Earlier work this paper cites.
arXiv preprint arXiv:1809.11096 (2018)
Brock, A., Donahue, J., Simonyan, K.: Large scale gan training for high fidelity natural image synthesis · 2018
Earlier work this paper cites.
In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1316–1324 (2018)
Xu, T., Zhang, P., Huang, Q., Zhang, H., Gan, Z., Huang, X., He, X.: Attngan: Fine-grained text to image generation with attentional generative adversarial networks · 2018
Earlier work this paper cites.
IEEE transactions on pattern analysis and machine intelligence 41
Zhang, H., Xu, T., Li, H., Zhang, S., Wang, X., Huang, X., Metaxas, D.N.: Stackgan++: Realistic image synthesis with stacked generative adversarial networks · 2018
Earlier work this paper cites.
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10,051–10,060 (2019)
Kolkin, N., Salavon, J., Shakhnarovich, G.: Style transfer by relaxed optimal transport and self-similarity · 2019
Earlier work this paper cites.
Advances in neural information processing systems 32
Li, B., Qi, X., Lukasiewicz, T., Torr, P.: Controllable text-to-image generation · 2019
Earlier work this paper cites.
Advances in neural information processing systems 33
Ho, J., Jain, A., Abbeel, P.: Denoising diffusion probabilistic models · 2020
Earlier work this paper cites.
https://github.com/mseitzer/pytorch-fid
Seitzer, M.: pytorch-fid: FID Score for PyTorch · 2020
Earlier work this paper cites.
arXiv preprint arXiv:2010.02502 (2020)
Song, J., Meng, C., Ermon, S.: Denoising diffusion implicit models · 2020
Earlier work this paper cites.
Advances in neural information processing systems 34
Dhariwal, P., Nichol, A.: Diffusion models beat gans on image synthesis · 2021
Earlier work this paper cites.
Advances in neural information processing systems 34
Ding, M., Yang, Z., Hong, W., Zheng, W., Zhou, C., Yin, D., Lin, J., Zou, X., Shao, Z., Yang, H., et al.: Cogview: Mastering text-to-image generation via transformers · 2021
Earlier work this paper cites.
In: International Conference on Learning Representations (2021)
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al.: An image is worth 16x16 words: Transformers for image recognition at scale · 2021
Cited alongside, same era.
arXiv preprint arXiv:2106.09685 (2021)
Hu, E.J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W.: Lora: Low-rank adaptation of large language models · 2021
Cited alongside, same era.
In: International conference on machine learning, pp. 8162–8171. PMLR (2021)
Nichol, A.Q., Dhariwal, P.: Improved denoising diffusion probabilistic models · 2021
Cited alongside, same era.
In: International conference on machine learning, pp. 8748–8763. PMLR (2021)
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision · 2021
Cited alongside, same era.
In: International conference on machine learning, pp. 8821–8831. Pmlr (2021)
Advances in Neural Information Processing Systems (2022)
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al.: Laion-5b: An open large-scale dataset for training next generation image-text models · 2022
Later among the works it cites.
arXiv preprint arXiv:2211.13203 (2022)
Zhang, Y., Huang, N., Tang, F., Huang, H., Ma, C., Dong, W., Xu, C.: Inversion-based creativity transfer with diffusion models · 2022
Later among the works it cites.
In: ACM SIGGRAPH 2022 conference proceedings, pp. 1–8 (2022)
Zhang, Y., Tang, F., Dong, W., Huang, H., Ma, C., Lee, T.Y., Xu, C.: Domain enhanced arbitrary image style transfer via contrastive learning · 2022
Later among the works it cites.
Li, J., Li, D., Savarese, S., Hoi, S.: Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models pp. 19,730–19,742 (2023)
2023
Closest in time.
Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., Aberman, K.: Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation pp. 22,500–22,510 (2023)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., Sutskever, I.: Zero-shot text-to-image generation · 2021
Cited alongside, same era.
In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 14,618–14,627 (2021)
Wu, X., Hu, Z., Sheng, L., Xu, D.: Styleformer: Real-time arbitrary style transfer via parametric style composition · 2021
Cited alongside, same era.
arXiv preprint arXiv:2111.13792 (2021)
Zhou, Y., Zhang, R., Chen, C., Li, C., Tensmeyer, C., Yu, T., Gu, J., Xu, J., Sun, T.: Lafite: Towards language-free training for text-to-image generation · 2021
Cited alongside, same era.
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 11,326–11,336 (2022)
Deng, Y., Tang, F., Dong, W., Ma, C., Pan, X., Wang, L., Xu, C.: Stytr2: Image style transfer with transformers · 2022
Cited alongside, same era.
In: European Conference on Computer Vision, pp. 89–106. Springer (2022)
Gafni, O., Polyak, A., Ashual, O., Sheynin, S., Parikh, D., Taigman, Y.: Make-a-scene: Scene-based text-to-image generation with human priors · 2022
Cited alongside, same era.
arXiv preprint arXiv:2208.01618 (2022)
Gal, R., Alaluf, Y., Atzmon, Y., Patashnik, O., Bermano, A.H., Chechik, G., Cohen-Or, D.: An image is worth one word: Personalizing text-to-image generation using textual inversion · 2022
Cited alongside, same era.
Journal of Machine Learning Research 23
Ho, J., Saharia, C., Chan, W., Fleet, D.J., Norouzi, M., Salimans, T.: Cascaded diffusion models for high fidelity image generation · 2022
Cited alongside, same era.
In: International Conference on Machine Learning (2022)
Nichol, A.Q., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., Mcgrew, B., Sutskever, I., Chen, M.: Glide: Towards photorealistic image generation and editing with text-guided diffusion models · 2022
Cited alongside, same era.
2023
Closest in time.
Advances in Neural Information Processing Systems (2023)
Sohn, K., Ruiz, N., Lee, K., Chin, D.C., Blok, I., Chang, H., Barber, J., Jiang, L., Entis, G., Li, Y., et al.: Styledrop: Text-to-image generation in any style · 2023
Closest in time.
arXiv preprint arXiv:2303.09522 (2023)
Voynov, A., Chu, Q., Cohen-Or, D., Aberman, K.: p + p+ : Extended textual conditioning in text-to-image generation · 2023
Closest in time.
arXiv preprint arXiv:2308.06721 (2023)
Ye, H., Zhang, J., Liu, S., Han, X., Yang, W.: Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models · 2023
Closest in time.
Zhang, L., Rao, A., Agrawala, M.: Adding conditional control to text-to-image diffusion models pp. 3836–3847 (2023)
2023
Closest in time.
ACM Transactions on Graphics (TOG) 42
Zhang, Y., Dong, W., Tang, F., Huang, N., Huang, H., Ma, C., Lee, T.Y., Deussen, O., Xu, C.: Prospect: Prompt spectrum for attribute-aware personalization of diffusion models · 2023
Closest in time.
Zhang, Y., Huang, N., Tang, F., Huang, H., Ma, C., Dong, W., Xu, C.: Inversion-based style transfer with diffusion models pp. 10,146–10,156 (2023)
2023
Closest in time.
International Journal of Computer Vision pp. 1–16 (2024)
Chen, T., Pu, T., Liu, L., Shi, Y., Yang, Z., Lin, L.: Heterogeneous semantic transfer for multi-label recognition with partial labels · 2024
Closest in time.
IEEE Transactions on Image Processing (2024)
Chen, T., Wang, W., Pu, T., Qin, J., Yang, Z., Liu, J., Lin, L.: Dynamic correlation learning and regularization for multi-label confidence calibration · 2024
Closest in time.
Mou, C., Wang, X., Xie, L., Wu, Y., Zhang, J., Qi, Z., Shan, Y.: T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models 38
2024
Closest in time.
International Conference on Learning Representations (2024)
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., Rombach, R.: Sdxl: Improving latent diffusion models for high-resolution image synthesis · 2024
Closest in time.
arXiv preprint arXiv:2404.02733 (2024)
Wang, H., Wang, Q., Bai, X., Qin, Z., Chen, A.: Instantstyle: Free lunch towards style-preserving in text-to-image generation · 2024
Closest in time.