Fetching the paper…
Reading the bibliography…
Text-driven diffusion models have become increasingly popular for various image editing tasks, including inpainting, stylization, and object replacement.
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollar, P., Zitnick, L.: Microsoft coco: Common objects in context. In: ECCV. ECCV (September 2014), https://www.microsoft.com/en-us/research/publication/microsoft-coco-common-objects-in-context/
2014
Earlier work this paper cites.
Ronneberger, O., Fischer, P., Brox, T.: U-net: Convolutional networks for biomedical image segmentation. In: Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18. pp. 234–241. Springer (2015)
2015
Earlier work this paper cites.
Dong, C., Loy, C.C., He, K., Tang, X.: Image super-resolution using deep convolutional networks. IEEE Trans. Pattern Anal. Mach. Intell. 38
2016
Earlier work this paper cites.
Agustsson, E., Timofte, R.: NTIRE 2017 challenge on single image super-resolution: Dataset and study. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops, CVPR Workshops 2017, Honolulu, HI, USA, July 21-26, 2017. pp. 1122–1131. IEEE Computer Society (2017). https://doi.org/10.1109/CVPRW.2017.150, https://doi.org/10.1109/CVPRW.2017.150
2017
Earlier work this paper cites.
Galteri, L., Seidenari, L., Bertini, M., Bimbo, A.: Deep generative adversarial compression artifact removal. In: ICCV (2017)
2017
Earlier work this paper cites.
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., Hochreiter, S.: Gans trained by a two time-scale update rule converge to a local nash equilibrium. In: NeurIPS (2017)
2017
Earlier work this paper cites.
Lim, B., Son, S., Kim, H., Nah, S., Lee, K.M.: Enhanced deep residual networks for single image super-resolution. In: Proceedings of CVPR Workshops (2017)
2017
Earlier work this paper cites.
Su, S., Delbracio, M., Wang, J., Sapiro, G., Heidrich, W., Wang, O.: Deep video deblurring for hand-held cameras. In: CVPR. pp. 1279–1288 (2017)
2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. NeurIPS 30
2017
Earlier work this paper cites.
Zhang, K., Zuo, W., Chen, Y., Meng, D., Zhang, L.: Beyond a Gaussian denoiser: Residual learning of deep CNN for image denoising. IEEE Transactions on Image Processing 26
2017
Earlier work this paper cites.
Blau, Y., Michaeli, T.: The perception-distortion tradeoff. In: CVPR. pp. 6228–6237 (2018)
2018
Earlier work this paper cites.
Wang, X., Yu, K., Wu, S., Gu, J., Liu, Y., Dong, C., Qiao, Y., Loy, C.C.: ESRGAN: enhanced super-resolution generative adversarial networks. In: Proceedings of ECCV Workshops (2018)
2018
Earlier work this paper cites.
Zhang, K., Zuo, W., Zhang, L.: Ffdnet: Toward a fast and flexible solution for CNN based image denoising. IEEE Transactions on Image Processing (2018)
2018
Earlier work this paper cites.
Zhang, R., Isola, P., Efros, A.A., Shechtman, E., Wang, O.: The unreasonable effectiveness of deep features as a perceptual metric. In: CVPR (2018)
2018
Earlier work this paper cites.
Nah, S., Timofte, R., Baik, S., Hong, S., Moon, G., Son, S., Lee, K.M., Wang, X., Chan, K.C.K., Yu, K., Dong, C., Loy, C.C., Fan, Y., Yu, J., Liu, D., Huang, T.S., Sim, H., Kim, M., Park, D., Kim, J., Chun, S.Y., Haris, M., Shakhnarovich, G., Ukita, N., Zamir, S.W., Arora, A., Khan, S.H., Khan, F.S., Shao, L., Gupta, R.K., Chudasama, V.M., Patel, H., Upla, K.P., Fan, H., Li, G., Zhang, Y., Li, X., Zhang, W., He, Q., Purohit, K., Rajagopalan, A.N., Kim, J., Tofighi, M., Guo, T., Monga, V.: NTIRE 2019 challenge on video deblurring: Methods and results. In: IEEE Conference on Computer Vision and Pattern Recognition Workshops, CVPR Workshops 2019, Long Beach, CA, USA, June 16-20, 2019. pp. 1974–1984. Computer Vision Foundation / IEEE (2019). https://doi.org/10.1109/CVPRW.2019.00249, http://openaccess.thecvf.com/content_CVPRW_2019/html/NTIRE/Nah_NTIRE_2019_Challenge_on_Video_Deblurring_Methods_and_Results_CVPRW_2019_paper.html
2019
Earlier work this paper cites.
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al.: Language models are unsupervised multitask learners. OpenAI blog 1
2019
Earlier work this paper cites.
Song, Y., Ermon, S.: Generative modeling by estimating gradients of the data distribution. In: NeurIPS. pp. 11895–11907 (2019)
2019
Earlier work this paper cites.
Ho, J., Jain, A., Abbeel, P.: Denoising diffusion probabilistic models. NeurIPS 33
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Abuolaim, A., Delbracio, M., Kelly, D., Brown, M.S., Milanfar, P.: Learning to reduce defocus blur by realistically modeling dual-pixel data. In: ICCV. pp. 2289–2298 (2021)
2021
Earlier work this paper cites.
Delbracio, M., Talebei, H., Milanfar, P.: Projected distribution loss for image enhancement. In: 2021 IEEE International Conference on Computational Photography (ICCP). pp. 1–12. IEEE (2021)
2021
Earlier work this paper cites.
Dhariwal, P., Nichol, A.: Diffusion models beat gans on image synthesis. Neural Information Processing Systems (2021)
2021
Earlier work this paper cites.
Gu, X., Lin, T.Y., Kuo, W., Cui, Y.: Open-vocabulary object detection via vision and language knowledge distillation. In: ICLR (2021)
2021
Earlier work this paper cites.
Ho, J., Salimans, T.: Classifier-free diffusion guidance. In: NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications (2021), https://openreview.net/forum?id=qw8AKxfYbI
2021
Earlier work this paper cites.
Ke, J., Wang, Q., Wang, Y., Milanfar, P., Yang, F.: Musiq: Multi-scale image quality transformer. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 5148–5157 (2021)
2021
Earlier work this paper cites.
Liang, J., Cao, J., Sun, G., Zhang, K., Gool, L.V., Timofte, R.: Swinir: Image restoration using swin transformer. In: Proceedings of ICCV Workshops (2021)
2021
Earlier work this paper cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: International conference on machine learning. pp. 8748–8763. PMLR (2021)
2021
Earlier work this paper cites.
Song, Y., Sohl-Dickstein, J., Kingma, D.P., Kumar, A., Ermon, S., Poole, B.: Score-based generative modeling through stochastic differential equations. In: ICLR (2021), https://openreview.net/forum?id=PxTIG12RRHS
2021
Earlier work this paper cites.
Wang, L., Wang, Y., Lin, Z., Yang, J., An, W., Guo, Y.: Learning a single network for scale-arbitrary super-resolution. In: ICCV. pp. 4801–4810 (2021)
2021
Earlier work this paper cites.
Wang, X., Li, Y., Zhang, H., Shan, Y.: Towards real-world blind face restoration with generative facial prior. In: CVPR (2021)
2021
Cited alongside, same era.
Wang, X., Xie, L., Dong, C., Shan, Y.: Real-esrgan: Training real-world blind super-resolution with pure synthetic data. In: ICCV. pp. 1905–1914 (2021)
2021
Cited alongside, same era.
Yang, T., Ren, P., Xie, X., Zhang, L.: Gan prior embedded network for blind face restoration in the wild. In: CVPR (2021)
2021
Cited alongside, same era.
Zamir, S.W., Arora, A., Khan, S., Hayat, M., Khan, F.S., Yang, M.H., Shao, L.: Multi-stage progressive image restoration. In: CVPR (2021)
2021
Cited alongside, same era.
Zhang, K., Liang, J., Van Gool, L., Timofte, R.: Designing a practical degradation model for deep blind image super-resolution. In: ICCV. pp. 4791–4800 (2021)
2021
Cited alongside, same era.
2023
Closest in time.
Delbracio, M., Milanfar, P.: Inversion by direct iteration: An alternative to denoising diffusion for image restoration. Transactions on Machine Learning Research (2023), https://openreview.net/forum?id=VmyFF5lL3F , featured Certification
2023
Closest in time.
Gal, R., Alaluf, Y., Atzmon, Y., Patashnik, O., Bermano, A.H., Chechik, G., Cohen-Or, D.: An image is worth one word: Personalizing text-to-image generation using textual inversion. In: ICLR (2023)
2023
Closest in time.
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Avrahami, O., Lischinski, D., Fried, O.: Blended diffusion for text-driven editing of natural images. In: CVPR. pp. 18208–18218 (2022)
2022
Cited alongside, same era.
Couairon, G., Verbeek, J., Schwenk, H., Cord, M.: Diffedit: Diffusion-based semantic image editing with mask guidance. In: ICLR (2022)
2022
Cited alongside, same era.
Gu, J., Cai, H., Dong, C., Ren, J.S., Timofte, R., Gong, Y., Lao, S., Shi, S., Wang, J., Yang, S., Wu, T., Xia, W., Yang, Y., Cao, M., Heng, C., Fu, L., Zhang, R., Zhang, Y., Wang, H., Song, H., Wang, J., Fan, H., Hou, X., Sun, M., Li, M., Zhao, K., Yuan, K., Kong, Z., Wu, M., Zheng, C., Conde, M.V., Burchi, M., Feng, L., Zhang, T., Li, Y., Xu, J., Wang, H., Liao, Y., Li, J., Xu, K., Sun, T., Xiong, Y., Keshari, A., Komal, Thakur, S., Jakhetiya, V., Subudhi, B.N., Yang, H.H., Chang, H.E., Huang, Z.K., Chen, W.T., Kuo, S.Y., Dutta, S., Das, S.D., Shah, N.A., Tiwari, A.K.: Ntire 2022 challenge on perceptual image quality assessment. In: CVPRW. pp. 951–967 (June 2022)
2022
Cited alongside, same era.
Hu, E.J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W.: LoRA: Low-rank adaptation of large language models. In: ICLR (2022), https://openreview.net/forum?id=nZeVKeeFYf9
2022
Cited alongside, same era.
Meng, C., He, Y., Song, Y., Song, J., Wu, J., Zhu, J.Y., Ermon, S.: SDEdit: Guided image synthesis and editing with stochastic differential equations. In: ICLR (2022)
2022
Cited alongside, same era.
Paiss, R., Chefer, H., Wolf, L.: No token left behind: Explainability-aided image classification and generation. In: ECCV (2022)
2022
Cited alongside, same era.
Prakash, M., Delbracio, M., Milanfar, P., Jug, F.: Interpretable unsupervised diversity denoising and artefact removal. In: ICLR (2022), https://openreview.net/forum?id=DfMqlB0PXjM
2022
Cited alongside, same era.
Han, L., Li, Y., Zhang, H., Milanfar, P., Metaxas, D., Yang, F.: Svdiff: Compact parameter space for diffusion fine-tuning. In: ICCV. pp. 7323–7334 (October 2023)
2023
Closest in time.
Hertz, A., Mokady, R., Tenenbaum, J., Aberman, K., Pritch, Y., Cohen-or, D.: Prompt-to-prompt image editing with cross-attention control. In: ICLR (2023), https://openreview.net/forum?id=_CDixzkzeyb
2023
Closest in time.
2023
Closest in time.
Ke, J., Ye, K., Yu, J., Wu, Y., Milanfar, P., Yang, F.: Vila: Learning image aesthetics from user comments with vision-language pretraining. In: CVPR. pp. 10041–10051 (2023)
2023
Closest in time.
Kumari, N., Zhang, B., Zhang, R., Shechtman, E., Zhu, J.Y.: Multi-concept customization of text-to-image diffusion. In: CVPR. pp. 1931–1941 (2023)
2023
Closest in time.
Liang, Z., Li, C., Zhou, S., Feng, R., Loy, C.C.: Iterative prompt learning for unsupervised backlit image enhancement. In: ICCV. pp. 8094–8103 (2023)
2023
Closest in time.
2023
Closest in time.
Liu, H., Li, C., Wu, Q., Lee, Y.J.: Visual instruction tuning. In: NeurIPS (2023)
2023
Closest in time.
Mokady, R., Hertz, A., Aberman, K., Pritch, Y., Cohen-Or, D.: Null-text inversion for editing real images using guided diffusion models. In: CVPR. pp. 6038–6047 (June 2023)
2023
Closest in time.
2023
Closest in time.
OpenAI: Gpt-4 technical report (2023)
2023
Closest in time.
Paiss, R., Ephrat, A., Tov, O., Zada, S., Mosseri, I., Irani, M., Dekel, T.: Teaching clip to count to ten. In: ICCV (2023)
2023
Closest in time.
Parmar, G., Kumar Singh, K., Zhang, R., Li, Y., Lu, J., Zhu, J.Y.: Zero-shot image-to-image translation. In: ACM SIGGRAPH 2023 Conference Proceedings. pp. 1–11 (2023)
2023
Closest in time.
Ren, M., Delbracio, M., Talebi, H., Gerig, G., Milanfar, P.: Multiscale structure guided diffusion for image deblurring. In: ICCV. pp. 10721–10733 (2023)
2023
Closest in time.
2023
Closest in time.
Tumanyan, N., Geyer, M., Bagon, S., Dekel, T.: Plug-and-play diffusion features for text-driven image-to-image translation. In: CVPR. pp. 1921–1930 (2023)
2023
Closest in time.
Wang, J., Chan, K.C., Loy, C.C.: Exploring clip for assessing the look and feel of images. In: AAAI (2023)
2023
Closest in time.
2023
Closest in time.
Wang, X., Chen, X., Ni, B., Wang, H., Tong, Z., Liu, Y.: Deep arbitrary-scale image super-resolution via scale-equivariance pursuit. In: CVPR (2023)
2023
Closest in time.
Zhang, L., Rao, A., Agrawala, M.: Adding conditional control to text-to-image diffusion models. In: ICCV (2023)
2023
Closest in time.
Luo, Z., Gustafsson, F.K., Zhao, Z., Sjölund, J., Schön, T.B.: Controlling vision-language models for multi-task image restoration. In: The Twelfth International Conference on Learning Representations (2024), https://openreview.net/forum?id=t3vnnLeajU
2024
Closest in time.
Wu, R., Yang, T., Sun, L., Zhang, Z., Li, S., Zhang, L.: Seesr: Towards semantics-aware real-world image super-resolution. In: CVPR (2024)
2024
Closest in time.
Yu, F., Gu, J., Li, Z., Hu, J., Kong, X., Wang, X., He, J., Qiao, Y., Dong, C.: Scaling up to excellence: Practicing model scaling for photo-realistic image restoration in the wild. In: CVPR (2024)
2024
Closest in time.
Zhang, K., Mo, L., Chen, W., Sun, H., Su, Y.: Magicbrush: A manually annotated dataset for instruction-guided image editing. Advances in Neural Information Processing Systems 36
2024
Closest in time.