Fetching the paper…
Reading the bibliography…
While large-scale pre-trained text-to-image models can synthesize diverse and high-quality human-centered images, novel challenges arise with a nuanced task of "identity fine editing": precisely modifying specific features of a subject while maintaining its inherent identity and context.
Digital image processing
R. C. Gonzales and P. Wintz · 1987
Earlier work this paper cites.
Tokenization as the initial phase in nlp
J. J. Webster and C. Kit · 1992
Earlier work this paper cites.
Face recognition by humans: Nineteen results all computer vision researchers should know about
P. Sinha, B. Balas, Y. Ostrovsky, and R. Russell · 2006
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Deepface: Closing the gap to human-level performance in face verification
Y. Taigman, M. Yang, M. Ranzato, and L. Wolf · 2014
Earlier work this paper cites.
Perceptual losses for real-time style transfer and super-resolution
J. Johnson, A. Alahi, and L. Fei-Fei · 2016
Earlier work this paper cites.
Progressive growing of gans for improved quality, stability, and variation
T. Karras, T. Aila, S. Laine, and J. Lehtinen · 2017
Earlier work this paper cites.
Photorealistic facial texture inference using deep neural networks
S. Saito, L. Wei, L. Hu, K. Nagano, and H. Li · 2017
Earlier work this paper cites.
Arcface: Additive angular margin loss for deep face recognition
J. Deng, J. Guo, N. Xue, and S. Zafeiriou · 2018
Earlier work this paper cites.
Controllable text-to-image generation
B. Li, X. Qi, T. Lukasiewicz, and P. Torr · 2019
Earlier work this paper cites.
Rifegan: Rich feature generation for text-to-image synthesis from prior knowledge
J. Cheng, F. Wu, Y. Tian, L. Wang, and D. Tao · 2020
Earlier work this paper cites.
Ganspace: Discovering interpretable gan controls
E. Härkönen, A. Hertzmann, J. Lehtinen, and S. Paris · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2020
Earlier work this paper cites.
Df-gan: Deep fusion generative adversarial networks for text-to-image synthesis
M. Tao, H. Tang, S. Wu, N. Sebe, X.-Y. Jing, F. Wu, and B. Bao · 2020
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin · 2021
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
P. Esser, R. Rombach, and B. Ommer · 2021
Earlier work this paper cites.
Stylegan-nada: Clip-guided domain adaptation of image generators
R. Gal, O. Patashnik, H. Maron, G. Chechik, and D. Cohen-Or · 2021
Earlier work this paper cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
A. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen · 2021
Cited alongside, same era.
Styleclip: Text-driven manipulation of stylegan imagery
O. Patashnik, Z. Wu, E. Shechtman, D. Cohen-Or, and D. Lischinski · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Cited alongside, same era.
Dae-gan: Dynamic aspect-aware gan for text-to-image synthesis
S. Ruan, Y. Zhang, K. Zhang, Y. Fan, F. Tang, Q. Liu, and E. Chen · 2021
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. G. Lopes, B. K. Ayan, T. Salimans, et al · 2022
Later among the works it cites.
Scaling autoregressive models for content-rich text-to-image generation
J. Yu, Y. Xu, J. Y. Koh, T. Luong, G. Baid, Z. Wang, V. Vasudevan, A. Ku, Y. Yang, B. K. Ayan, et al · 2022
Later among the works it cites.
Blended latent diffusion
O. Avrahami, O. Fried, and D. Lischinski · 2023
Later among the works it cites.
Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing
M. Cao, X. Wang, Z. Qi, Y. Shan, X. Qie, and Y. Zheng · 2023
Later among the works it cites.
Muse: Text-to-image generation via masked generative transformers
H. Chang, H. Zhang, J. Barber, A. Maschinot, J. Lezama, L. Jiang, M.-H. Yang, K. Murphy, W. T. Freeman, M. Rubinstein, et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Zhan, Y. Yu, R. Wu, J. Zhang, S. Lu, L. Liu, A. Kortylewski, C. Theobalt, and E. Xing · 2021
Cited alongside, same era.
ediffi: Text-to-image diffusion models with an ensemble of expert denoisers
Y. Balaji, S. Nah, X. Huang, A. Vahdat, J. Song, K. Kreis, M. Aittala, T. Aila, S. Laine, B. Catanzaro, et al · 2022
Cited alongside, same era.
Dreamartist: Towards controllable one-shot text-to-image generation via contrastive prompt-tuning
Z. Dong, P. Wei, and L. Lin · 2022
Cited alongside, same era.
An image is worth one word: Personalizing text-to-image generation using textual inversion
R. Gal, Y. Alaluf, Y. Atzmon, O. Patashnik, A. H. Bermano, G. Chechik, and D. Cohen-Or · 2022
Cited alongside, same era.
Prompt-to-prompt image editing with cross attention control
A. Hertz, R. Mokady, J. Tenenbaum, K. Aberman, Y. Pritch, and D. Cohen-Or · 2022
Cited alongside, same era.
Elucidating the design space of diffusion-based generative models
T. Karras, M. Aittala, T. Aila, and S. Laine · 2022
Cited alongside, same era.
Diffusionclip: Text-guided diffusion models for robust image manipulation
G. Kim, T. Kwon, and J. C. Ye · 2022
Cited alongside, same era.
Later among the works it cites.
Dreamidentity: Improved editability for efficient face-identity preserved image generation
Z. Chen, S. Fang, W. Liu, Q. He, M. Huang, Y. Zhang, and Z. Mao · 2023
Later among the works it cites.
Pfb-diff: Progressive feature blending diffusion for text-driven image editing
W. Huang, S. Tu, and L. Xu · 2023
Later among the works it cites.
Unified multi-modal latent diffusion for joint subject and text conditional image generation
Y. Ma, H. Yang, W. Wang, J. Fu, and J. Liu · 2023
Later among the works it cites.
Null-text inversion for editing real images using guided diffusion models
R. Mokady, A. Hertz, K. Aberman, Y. Pritch, and D. Cohen-Or · 2023
Later among the works it cites.
Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models
N. Ruiz, Y. Li, V. Jampani, W. Wei, T. Hou, Y. Pritch, N. Wadhwa, M. Rubinstein, and K. Aberman · 2023
Later among the works it cites.
Instantbooth: Personalized text-to-image generation without test-time finetuning
J. Shi, W. Xiong, Z. Lin, and H. J. Jung · 2023
Later among the works it cites.
Ledits: Real image editing with ddpm inversion and semantic guidance
L. Tsaban and A. Passos · 2023
Later among the works it cites.
Plug-and-play diffusion features for text-driven image-to-image translation
N. Tumanyan, M. Geyer, S. Bagon, and T. Dekel · 2023
Later among the works it cites.
Edict: Exact diffusion inversion via coupled transformations
B. Wallace, A. Gokul, and N. Naik · 2023
Later among the works it cites.
Elite: Encoding visual concepts into textual embeddings for customized text-to-image generation
Y. Wei, Y. Zhang, Z. Ji, J. Bai, L. Zhang, and W. Zuo · 2023
Later among the works it cites.
Uncovering the disentanglement capability in text-to-image diffusion models
Q. Wu, Y. Liu, H. Zhao, A. Kale, T. Bui, T. Yu, Z. Lin, Y. Zhang, and S. Chang · 2023
Later among the works it cites.
Sine: Single image editing with text-to-image diffusion models
Z. Zhang, L. Han, A. Ghosh, D. N. Metaxas, and J. Ren · 2023
Later among the works it cites.