Fetching the paper…
Reading the bibliography…
Subject-driven image inpainting has recently gained prominence in image editing with the rapid advancement of diffusion models.
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pp. 740–755 (2014). Springer
2014
Earlier work this paper cites.
Ronneberger, O., Fischer, P., Brox, T.: U-net: Convolutional networks for biomedical image segmentation. In: Medical Image Computing and Computer-assisted intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18, pp. 234–241 (2015). Springer
2015
Earlier work this paper cites.
Ho, J., Jain, A., Abbeel, P.: Denoising diffusion probabilistic models. Advances in neural information processing systems 33
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al
2021
Earlier work this paper cites.
Avrahami, O., Lischinski, D., Fried, O.: Blended diffusion for text-driven editing of natural images. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 18208–18218 (2022)
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Ho, J., Salimans, T.: Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598 (2022)
2022
Earlier work this paper cites.
Lugmayr, A., Danelljan, M., Romero, A., Yu, F., Timofte, R., Van Gool, L.: Repaint: Inpainting using denoising diffusion probabilistic models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11461–11471 (2022)
2022
Earlier work this paper cites.
Navasardyan, S., Ohanyan, M.: The family of onion convolutions for image inpainting. International Journal of Computer Vision 130
2022
Earlier work this paper cites.
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10684–10695 (2022)
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E.L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al
2022
Earlier work this paper cites.
Platen, P., Patil, S., Lozhkov, A., Cuenca, P., Lambert, N., Rasul, K., Davaadorj, M., Nair, D., Paul, S., Berman, W., Xu, Y., Liu, S., Wolf, T.: Diffusers: State-of-the-art diffusion models. GitHub (2022)
2022
Earlier work this paper cites.
Avrahami, O., Fried, O., Lischinski, D.: Blended latent diffusion. ACM transactions on graphics (TOG) 42
2023
Earlier work this paper cites.
Brooks, T., Holynski, A., Efros, A.A.: Instructpix2pix: Learning to follow image editing instructions. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 18392–18402 (2023)
2023
Earlier work this paper cites.
Diffusers: Stable-Diffusion-XL-1.0-Inpainting-0.1 Model Card, https://huggingface.co/diffusers/stable-diffusion-xl-1.0-inpainting-0.1 (2023)
2023
Earlier work this paper cites.
Huang, K., Sun, K., Xie, E., Li, Z., Liu, X.: T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation. Advances in Neural Information Processing Systems 36
2023
Earlier work this paper cites.
Kumari, N., Zhang, B., Zhang, R., Shechtman, E., Zhu, J.-Y.: Multi-concept customization of text-to-image diffusion. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1931–1941 (2023)
2023
Cited alongside, same era.
Labs, B.F.: FLUX.1 dev Model Card, https://huggingface.co/black-forest-labs/FLUX.1-dev (2023)
2023
Cited alongside, same era.
Labs, B.F.: FLUX.1 Fill dev Model Card, https://huggingface.co/black-forest-labs/FLUX.1-Fill-dev (2023)
2023
Cited alongside, same era.
Llyasviel: Fooocus Inpaint Model Card, https://huggingface.co/lllyasviel/fooocus_inpaint (2023)
2023
Cited alongside, same era.
liuhaotian: llava-v1.5-7b Model Card, https://huggingface.co/liuhaotian/llava-v1.5-7b (2024)
2024
Closest in time.
Li, P., Nie, Q., Chen, Y., Jiang, X., Wu, K., Lin, Y., Liu, Y., Peng, J., Wang, C., Zheng, F.: Tuning-free image customization with image and text guidance. In: European Conference on Computer Vision, pp. 233–250 (2024). Springer
2024
Closest in time.
OpenAI: ChatGPT-4o, https://openai.com/chatgpt/ (2024)
2024
Closest in time.
2024
Closest in time.
Quan, W., Chen, J., Liu, Y., Yan, D.-M., Wonka, P.: Deep learning-based image and video inpainting: A survey. International Journal of Computer Vision 132
2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Patashnik, O., Garibi, D., Azuri, I., Averbuch-Elor, H., Cohen-Or, D.: Localizing object-level shape variations with text-to-image diffusion models. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 23051–23061 (2023)
2023
Cited alongside, same era.
Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., Aberman, K.: Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 22500–22510 (2023)
2023
Cited alongside, same era.
RunwayML: Stable Diffusion Inpainting model card, https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-inpainting (2023)
2023
Cited alongside, same era.
Stabilityai: Stable Diffusion v2 Inpainting Model Card, https://huggingface.co/stabilityai/stable-diffusion-2-inpainting (2023)
2023
Cited alongside, same era.
Tumanyan, N., Geyer, M., Bagon, S., Dekel, T.: Plug-and-play diffusion features for text-driven image-to-image translation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1921–1930 (2023)
2023
Cited alongside, same era.
Wei, Y., Zhang, Y., Ji, Z., Bai, J., Zhang, L., Zuo, W.: Elite: Encoding visual concepts into textual embeddings for customized text-to-image generation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 15943–15953 (2023)
2023
Cited alongside, same era.
Closest in time.
Qwen: Qwen2-VL-2B-Instruct Model Card, https://huggingface.co/Qwen/Qwen2-VL-2B-Instruct (2024)
2024
Closest in time.
Shi, J., Xiong, W., Lin, Z., Jung, H.J.: Instantbooth: Personalized text-to-image generation without test-time finetuning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8543–8552 (2024)
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Yang, N., Luan, X., Jia, H., Han, Z., Li, X., Tang, Y.: Ccr: Facial image editing with continuity, consistency and reversibility. International Journal of Computer Vision 132
2024
Closest in time.
2024
Closest in time.
Ge, S., Park, T., Zhu, J.-Y., Huang, J.-B.: Expressive image generation and editing with rich text. International Journal of Computer Vision 133
2025
Closest in time.
HuggingFaceTB: SmolVLM-500M-Instruct Model Card, https://huggingface.co/HuggingFaceTB/SmolVLM-500M-Instruct (2025)
2025
Closest in time.
Jin, J., Shen, Y., Zhao, X., Fu, Z., Yang, J.: Unicanvas: Affordance-aware unified real image editing via customized text-to-image generation. International Journal of Computer Vision 133
2025
Closest in time.
Li, X., Ren, Y., Jin, X., Lan, C., Wang, X., Zeng, W., Wang, X., Chen, Z.: Diffusion models for image restoration and enhancement: A comprehensive survey. International Journal of Computer Vision (2025)
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Song, D., Zhang, X., Zhou, J., Nie, W., Tong, R., Kankanhalli, M., Liu, A.-A.: Image-based virtual try-on: A survey. International Journal of Computer Vision 133
2025
Closest in time.
Wang, Q., Li, B., Li, X., Cao, B., Ma, L., Lu, H., Jia, X.: Characterfactory: Sampling consistent characters with gans for diffusion models. IEEE Transactions on Image Processing 34
2025
Closest in time.
Xing, J., Xu, C., Qian, Y., Liu, Y., Dai, G., Sun, B., Liu, Y., Wang, J.: Tryon-adapter: Efficient fine-grained clothing identity adaptation for high-fidelity virtual try-on. International Journal of Computer Vision 133
2025
Closest in time.
Yildirim, A.B., Pehlivan, H., Dundar, A.: Warping the residuals for image editing with stylegan. International Journal of Computer Vision 133
2025
Closest in time.