Fetching the paper…
Reading the bibliography…
A significant research effort is focused on exploiting the amazing capacities of pretrained diffusion models for the editing of images.They either finetune the model, or invert the image in the latent space of the pretrained model.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Earlier work this paper cites.
Precise recovery of latent vectors from generative adversarial networks
Z. C. Lipton and S. Tripathi · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Semantic image inpainting with deep generative models, 2017
R. A. Yeh, C. Chen, T. Y. Lim, A. G. Schwing, M. Hasegawa-Johnson, and M. N. Do · 2017
Earlier work this paper cites.
Inverting the generator of a generative adversarial network
A. Creswell and A. A. Bharath · 2018
Earlier work this paper cites.
Memory replay gans: Learning to generate new categories without forgetting
C. Wu, L. Herranz, X. Liu, J. Van De Weijer, B. Raducanu, et al · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang · 2018
Earlier work this paper cites.
Image2stylegan: How to embed images into the stylegan latent space?
R. Abdal, Y. Qin, and P. Wonka · 2019
Earlier work this paper cites.
Ganalyze: Toward visual definitions of cognitive image properties
L. Goetschalckx, A. Andonian, A. Oliva, and P. Isola · 2019
Earlier work this paper cites.
Image2stylegan++: How to edit the embedded images?
R. Abdal, Y. Qin, and P. Wonka · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
On the”steerability” of generative adversarial networks
A. Jahanian, L. Chai, and P. Isola · 2020
Earlier work this paper cites.
Encoding in style: a stylegan encoder for image-to-image translation
E. Richardson, Y. Alaluf, O. Patashnik, Y. Nitzan, Y. Azar, S. Shapiro, and D. Cohen-Or · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2020
Earlier work this paper cites.
In-domain gan inversion for real image editing
J. Zhu, Y. Shen, D. Zhao, and B. Zhou · 2020
Earlier work this paper cites.
Hyperstyle: Stylegan inversion with hypernetworks for real image editing, 2021
Y. Alaluf, O. Tov, R. Mokady, R. Gal, and A. H. Bermano · 2021
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
P. Dhariwal and A. Nichol · 2021
Earlier work this paper cites.
Stylegan-nada: Clip-guided domain adaptation of image generators
R. Gal, O. Patashnik, H. Maron, G. Chechik, and D. Cohen-Or · 2021
Earlier work this paper cites.
CLIPScore: a reference-free evaluation metric for image captioning
J. Hessel, A. Holtzman, M. Forbes, R. L. Bras, and Y. Choi · 2021
Earlier work this paper cites.
Classifier-free diffusion guidance
J. Ho and T. Salimans · 2021
Earlier work this paper cites.
Diff-tts: A denoising diffusion model for text-to-speech
M. Jeong, H. Kim, S. J. Cheon, B. J. Choi, and N. S. Kim · 2021
Earlier work this paper cites.
Sdedit: Image synthesis and editing with stochastic differential equations
C. Meng, Y. Song, J. Song, J. Wu, J.-Y. Zhu, and S. Ermon · 2021
Earlier work this paper cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
A. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen · 2021
Earlier work this paper cites.
Styleclip: Text-driven manipulation of stylegan imagery
O. Patashnik, Z. Wu, E. Shechtman, D. Cohen-Or, and D. Lischinski · 2021
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models, 2021
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2021
Earlier work this paper cites.
Designing an encoder for stylegan image manipulation
O. Tov, Y. Alaluf, Y. Nitzan, O. Patashnik, and D. Cohen-Or · 2021
Cited alongside, same era.
Gan inversion: A survey, 2021
W. Xia, Y. Zhang, Y. Yang, J.-H. Xue, B. Zhou, and M.-H. Yang · 2021
Cited alongside, same era.
O. Avrahami, O. Fried, and D. Lischinski · 2022
Cited alongside, same era.
Blended diffusion for text-driven editing of natural images
O. Avrahami, D. Lischinski, and O. Fried · 2022
Cited alongside, same era.
Diffedit: Diffusion-based semantic image editing with mask guidance
G. Couairon, J. Verbeek, H. Schwenk, and M. Cord · 2022
Cited alongside, same era.
Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing
M. Cao, X. Wang, Z. Qi, Y. Shan, X. Qie, and Y. Zheng · 2023
Closest in time.
Training-free layout control with cross-attention guidance
M. Chen, I. Laina, and A. Vedaldi · 2023
Closest in time.
Prompt tuning inversion for text-driven image editing using diffusion models
W. Dong, S. Xue, X. Duan, and S. Han · 2023
Closest in time.
Encoder-based domain tuning for fast personalization of text-to-image models
R. Gal, M. Arar, Y. Atzmon, A. H. Bermano, G. Chechik, and D. Cohen-Or · 2023
Closest in time.
Highly personalized text embedding for image manipulation by stable diffusion
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
O. Gafni, A. Polyak, O. Ashual, S. Sheynin, D. Parikh, and Y. Taigman · 2022
Cited alongside, same era.
An image is worth one word: Personalizing text-to-image generation using textual inversion
R. Gal, Y. Alaluf, Y. Atzmon, O. Patashnik, A. H. Bermano, G. Chechik, and D. Cohen-Or · 2022
Cited alongside, same era.
Prompt-to-prompt image editing with cross attention control
A. Hertz, R. Mokady, J. Tenenbaum, K. Aberman, Y. Pritch, and D. Cohen-Or · 2022
Cited alongside, same era.
Fastdiff: A fast conditional diffusion model for high-quality speech synthesis
R. Huang, M. W. Lam, J. Wang, D. Su, D. Yu, Y. Ren, and Z. Zhao · 2022
Cited alongside, same era.
Imagic: Text-based real image editing with diffusion models
B. Kawar, S. Zada, O. Lang, O. Tov, H.-T. Chang, T. Dekel, I. Mosseri, and M. Irani · 2022
Cited alongside, same era.
Diffusionclip: Text-guided diffusion models for robust image manipulation
G. Kim, T. Kwon, and J. C. Ye · 2022
Cited alongside, same era.
Specgrad: Diffusion probabilistic model based neural vocoder with adaptive noise spectral shaping
Y. Koizumi, H. Zen, K. Yatabe, N. Chen, and M. Bacchiani · 2022
Cited alongside, same era.
I. Han, S. Yang, T. Kwon, and J. C. Ye · 2023
Closest in time.
Improving negative-prompt inversion via proximal guidance
L. Han, S. Wen, Q. Chen, Z. Zhang, K. Song, M. Ren, R. Gao, Y. Chen, D. Liu, Q. Zhangli, et al · 2023
Closest in time.
Taming encoder for zero fine-tuning image customization with text-to-image diffusion models
X. Jia, Y. Zhao, K. C. Chan, Y. Li, H. Zhang, B. Gong, T. Hou, H. Wang, and Y.-C. Su · 2023
Closest in time.
Scaling up gans for text-to-image synthesis
M. Kang, J.-Y. Zhu, R. Zhang, J. Park, E. Shechtman, S. Paris, and T. Park · 2023
Closest in time.
Text2video-zero: Text-to-image diffusion models are zero-shot video generators
L. Khachatryan, A. Movsisyan, V. Tadevosyan, R. Henschel, Z. Wang, S. Navasardyan, and H. Shi · 2023
Closest in time.
Multi-concept customization of text-to-image diffusion
N. Kumari, B. Zhang, R. Zhang, E. Shechtman, and J.-Y. Zhu · 2023
Closest in time.
Differential diffusion: Giving each pixel its strength
E. Levin and O. Fried · 2023
Closest in time.
Magic3d: High-resolution text-to-3d content creation
C.-H. Lin, J. Gao, L. Tang, T. Takikawa, X. Zeng, X. Huang, K. Kreis, S. Fidler, M.-Y. Liu, and T.-Y. Lin · 2023
Closest in time.
Accelerating diffusion models for inverse problems through shortcut sampling
G. Liu, H. Sun, J. Li, F. Yin, and Y. Yang · 2023
Closest in time.
Negative-prompt inversion: Fast image inversion for editing with text-guided diffusion models
D. Miyake, A. Iohara, Y. Saito, and T. Tanaka · 2023
Closest in time.
Dragondiffusion: Enabling drag-style manipulation on diffusion models
C. Mou, X. Wang, J. Song, Y. Shan, and J. Zhang · 2023
Closest in time.
Zero-shot image-to-image translation
G. Parmar, K. K. Singh, R. Zhang, Y. Li, J. Lu, and J.-Y. Zhu · 2023
Closest in time.
Controlling text-to-image diffusion by orthogonal finetuning
Z. Qiu, W. Liu, H. Feng, Y. Xue, Y. Feng, Z. Liu, D. Zhang, A. Weller, and B. Schölkopf · 2023
Closest in time.
StyleGAN-T: Unlocking the power of GANs for fast large-scale text-to-image synthesis
A. Sauer, T. Karras, S. Laine, A. Geiger, and T. Aila · 2023
Closest in time.
Key-locked rank one editing for text-to-image personalization
Y. Tewel, R. Gal, G. Chechik, and Y. Atzmon · 2023
Closest in time.
Plug-and-play diffusion features for text-driven image-to-image translation
N. Tumanyan, M. Geyer, S. Bagon, and T. Dekel · 2023
Closest in time.
Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distillation
Z. Wang, C. Lu, Y. Wang, F. Bao, C. Li, H. Su, and J. Zhu · 2023
Closest in time.
Fastcomposer: Tuning-free multi-subject image generation with localized attention
G. Xiao, T. Yin, W. T. Freeman, F. Durand, and S. Han · 2023
Closest in time.
E. Xie, L. Yao, H. Shi, Z. Liu, D. Zhou, Z. Liu, J. Li, and Z. Li · 2023
Closest in time.
Adding conditional control to text-to-image diffusion models, 2023
L. Zhang and M. Agrawala · 2023
Closest in time.
Prospect: Expanded conditioning for the personalization of attribute-aware image generation
Y. Zhang, W. Dong, F. Tang, N. Huang, H. Huang, C. Ma, T.-Y. Lee, O. Deussen, and C. Xu · 2023
Closest in time.
Continuous layout editing of single images with diffusion models
Z. Zhang, Z. Huang, and J. Liao · 2023
Closest in time.