Fetching the paper…
Reading the bibliography…
Diffusion models continuously push the boundary of state-of-the-art image generation, but the process is hard to control with any nuance: practice proves that textual prompts are inadequate for accurately describing image style or fine structural details (such as faces).
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13. pp. 740–755. Springer (2014)
2014
Earlier work this paper cites.
Liu, Z., Luo, P., Wang, X., Tang, X.: Deep learning face attributes in the wild. In: Proceedings of International Conference on Computer Vision (ICCV) (December 2015)
2015
Earlier work this paper cites.
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., Ganguli, S.: Deep unsupervised learning using nonequilibrium thermodynamics. In: International conference on machine learning. pp. 2256–2265. PMLR (2015)
2015
Earlier work this paper cites.
Saleh, B., Elgammal, A.: Large-scale classification of fine-art paintings: Learning the right metric on the right feature. International Journal for Digital Art History (2) (2016)
2016
Earlier work this paper cites.
Zamir, A.R., Sax, A., Shen, W., Guibas, L.J., Malik, J., Savarese, S.: Taskonomy: Disentangling task transfer learning. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 3712–3722 (2018)
2018
Earlier work this paper cites.
Ho, J., Jain, A., Abbeel, P.: Denoising diffusion probabilistic models. Advances in neural information processing systems 33
2020
Earlier work this paper cites.
Song, J., Meng, C., Ermon, S.: Denoising diffusion implicit models. In: International Conference on Learning Representations (2020)
2020
Earlier work this paper cites.
Hessel, J., Holtzman, A., Forbes, M., Le Bras, R., Choi, Y.: Clipscore: A reference-free evaluation metric for image captioning. In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. pp. 7514–7528 (2021)
2021
Earlier work this paper cites.
Ho, J., Salimans, T.: Classifier-free diffusion guidance. In: NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications (2021)
2021
Earlier work this paper cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning transferable visual models from natural language supervision. In: ICML (2021)
2021
Earlier work this paper cites.
Hu, E.J., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al.: Lora: Low-rank adaptation of large language models. In: International Conference on Learning Representations (2022)
2022
Earlier work this paper cites.
Li, H., Yang, Y., Chang, M., Chen, S., Feng, H., Xu, Z., Li, Q., Chen, Y.: Srdiff: Single image super-resolution with diffusion probabilistic models. Neurocomputing 479
2022
Earlier work this paper cites.
Lugmayr, A., Danelljan, M., Romero, A., Yu, F., Timofte, R., Van Gool, L.: Repaint: Inpainting using denoising diffusion probabilistic models. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 11461–11471 (2022)
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 10684–10695 (2022)
2022
Earlier work this paper cites.
Saharia, C., Chan, W., Chang, H., Lee, C., Ho, J., Salimans, T., Fleet, D., Norouzi, M.: Palette: Image-to-image diffusion models. In: ACM SIGGRAPH 2022 conference proceedings. pp. 1–10 (2022)
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Cited alongside, same era.
Beaumont, R.: Vit-h/14 clip model (2023), https://huggingface.co/laion/CLIP-ViT-H-14-laion2B-s32B-b79K [Accessed: July 15th, 2024]
2024
Closest in time.
Conde, M.V., Geigle, G., Timofte, R.: High-quality image restoration following human instructions. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer (2024)
2024
Closest in time.
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., et al.: Scaling rectified flow transformers for high-resolution image synthesis. In: Forty-first International Conference on Machine Learning (2024)
2024
Closest in time.
Gokaslan, A., Cooper, A.F., Collins, J., Seguin, L., Jacobson, A., Patel, M., Frankle, J., Stephenson, C., Kuleshov, V.: Commoncanvas: Open diffusion models trained on creative-commons images. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 8250–8260 (2024)
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Brooks, T., Holynski, A., Efros, A.A.: Instructpix2pix: Learning to follow image editing instructions. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 18392–18402 (2023)
2023
Cited alongside, same era.
Hertz, A., Mokady, R., Tenenbaum, J., Aberman, K., Pritch, Y., Cohen-Or, D.: Prompt-to-prompt image editing with cross attention control. In: Proceedings of the International Conference on Learning Representations (ICLR) (2023)
2023
Cited alongside, same era.
Huang, L., Chen, D., Liu, Y., Shen, Y., Zhao, D., Zhou, J.: Composer: Creative and controllable image synthesis with composable conditions. In: International Conference on Machine Learning. pp. 13753–13773. PMLR (2023)
2023
Cited alongside, same era.
Pan, J., Sun, K., Ge, Y., Li, H., Duan, H., Wu, X., Zhang, R., Zhou, A., Qin, Z., Wang, Y., Dai, J., Qiao, Y., Li, H.: Journeydb: A benchmark for generative image understanding (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Xia, B., Zhang, Y., Wang, S., Wang, Y., Wu, X., Tian, Y., Yang, W., Van Gool, L.: Diffir: Efficient diffusion model for image restoration. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 13095–13105 (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Hao, Y., Chi, Z., Dong, L., Wei, F.: Optimizing prompts for text-to-image generation. Advances in Neural Information Processing Systems 36
2024
Closest in time.
Hristoforu, E.: Midjourneyimages dataset (2024), https://huggingface.co/datasets/ehristoforu/midjourney-images [Accessed: July 15th, 2024]
2024
Closest in time.
2024
Closest in time.
OpenAI: Gpt-4o (2024), https://openai.com/index/hello-gpt-4o/ [Accessed: July 15th, 2024]
2024
Closest in time.
Ruiz, N., Li, Y., Wadhwa, N., Pritch, Y., Rubinstein, M., Jacobs, D.E., Fruchter, S.: Magic insert: Style-aware drag-and-drop (2024)
2024
Closest in time.
https://huggingface.co/h94 : Pre-trained ipadapter+ weights for stable diffusion 1.5 (2024), https://huggingface.co/h94/IP-Adapter/blob/main/models/ip-adapter-plus_sd15.bin [Accessed: July 15th, 2024]
2024
Closest in time.
https://huggingface.co/h94 : Pre-trained scribble controlnet weights for stable diffusion 1.5 (2024), https://huggingface.co/lllyasviel/sd-controlnet-scribble [Accessed: July 15th, 2024]
2024
Closest in time.
https://huggingface.co/lllyasviel : Pre-trained openpose controlnet weights for stable diffusion 1.5 (2024), https://huggingface.co/lllyasviel/control_v11p_sd15_openpose [Accessed: July 15th, 2024]
2024
Closest in time.
2024
Closest in time.