Fetching the paper…
Reading the bibliography…
Large-scale text-to-image models including Stable Diffusion are capable of generating high-fidelity photorealistic portrait images.
Multi-concept customization of text-to-image diffusion
Kumari, N.; Zhang, B.; Zhang, R.; Shechtman, E.; and Zhu, J.-Y. 2023 · 1941
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J.; Meng, C.; and Ermon, S. 2020 · 2010
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Song, Y.; Sohl-Dickstein, J.; Kingma, D. P.; Kumar, A.; Ermon, S.; and Poole, B. 2020 · 2011
Earlier work this paper cites.
Image style transfer using convolutional neural networks
Gatys, L. A.; Ecker, A. S.; and Bethge, M. 2016 · 2016
Earlier work this paper cites.
Vggface2: A dataset for recognising faces across pose and age
Cao, Q.; Shen, L.; Xie, W.; Parkhi, O. M.; and Zisserman, A. 2018 · 2018
Earlier work this paper cites.
Arcface: Additive angular margin loss for deep face recognition
Deng, J.; Guo, J.; Xue, N.; and Zafeiriou, S. 2019 · 2019
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Song, Y.; and Ermon, S. 2019 · 2019
Earlier work this paper cites.
Few-shot adversarial learning of realistic neural talking head models
Zakharov, E.; Shysheya, A.; Burkov, E.; and Lempitsky, V. 2019 · 2019
Earlier work this paper cites.
Retinaface: Single-shot multi-level face localisation in the wild
Deng, J.; Guo, J.; Ververas, E.; Kotsia, I.; and Zafeiriou, S. 2020 · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Earlier work this paper cites.
Contrastive learning for unpaired image-to-image translation
Park, T.; Efros, A. A.; Zhang, R.; and Zhu, J.-Y. 2020 · 2020
Earlier work this paper cites.
In-domain gan inversion for real image editing
Zhu, J.; Shen, Y.; Zhao, D.; and Zhou, B. 2020 · 2020
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
Caron, M.; Touvron, H.; Misra, I.; Jégou, H.; Mairal, J.; Bojanowski, P.; and Joulin, A. 2021 · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Cited alongside, same era.
Noise2score: tweedie’s approach to self-supervised image denoising without clean images
Kim, K.; and Ye, J. C. 2021 · 2021
Cited alongside, same era.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Nichol, A.; Dhariwal, P.; Ramesh, A.; Shyam, P.; Mishkin, P.; McGrew, B.; Sutskever, I.; and Chen, M. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E. L.; Ghasemipour, K.; Gontijo Lopes, R.; Karagol Ayan, B.; Salimans, T.; et al. 2022 · 2022
Later among the works it cites.
Laion aesthetics
Schuhmann, C. Aug 2022 · 2022
Later among the works it cites.
Splicing vit features for semantic appearance transfer
Tumanyan, N.; Bar-Tal, O.; Bagon, S.; and Dekel, T. 2022 · 2022
Later among the works it cites.
Diffusers: State-of-the-art diffusion models
von Platen, P.; Patil, S.; Lozhkov, A.; Cuenca, P.; Lambert, N.; Rasul, K.; Davaadorj, M.; and Wolf, T. 2022 · 2022
Later among the works it cites.
Towards Robust Blind Face Restoration with Codebook Lookup TransFormer
Zhou, S.; Chan, K. C.; Li, C.; and Loy, C. C. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Real-ESRGAN: Training Real-World Blind Super-Resolution with Pure Synthetic Data
Wang, X.; Xie, L.; Dong, C.; and Shan, Y. 2021 · 2021
Cited alongside, same era.
An image is worth one word: Personalizing text-to-image generation using textual inversion
Gal, R.; Alaluf, Y.; Atzmon, Y.; Patashnik, O.; Bermano, A. H.; Chechik, G.; and Cohen-Or, D. 2022 · 2022
Cited alongside, same era.
Prompt-to-prompt image editing with cross attention control
Hertz, A.; Mokady, R.; Tenenbaum, J.; Aberman, K.; Pritch, Y.; and Cohen-Or, D. 2022 · 2022
Cited alongside, same era.
Diffusionclip: Text-guided diffusion models for robust image manipulation
Kim, G.; Kwon, T.; and Ye, J. C. 2022 · 2022
Cited alongside, same era.
Diffusion-based image translation using disentangled style and content representation
Kwon, G.; and Ye, J. C. 2022 · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022 · 2022
Cited alongside, same era.
Pivotal tuning for latent-based editing of real images
Roich, D.; Mokady, R.; Bermano, A. H.; and Cohen-Or, D. 2022 · 2022
Cited alongside, same era.
Later among the works it cites.
Break-A-Scene: Extracting Multiple Concepts from a Single Image
Avrahami, O.; Aberman, K.; Fried, O.; Cohen-Or, D.; and Lischinski, D. 2023 · 2023
Closest in time.
Svdiff: Compact parameter space for diffusion fine-tuning
Han, L.; Li, Y.; Zhang, H.; Milanfar, P.; Metaxas, D.; and Yang, F. 2023 · 2023
Closest in time.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Ruiz, N.; Li, Y.; Jampani, V.; Pritch, Y.; Rubinstein, M.; and Aberman, K. 2023 · 2023
Closest in time.
Instantbooth: Personalized text-to-image generation without test-time finetuning
Shi, J.; Xiong, W.; Lin, Z.; and Jung, H. J. 2023 · 2023
Closest in time.
Elite: Encoding visual concepts into textual embeddings for customized text-to-image generation
Wei, Y.; Zhang, Y.; Ji, Z.; Bai, J.; Zhang, L.; and Zuo, W. 2023 · 2023
Closest in time.
FastComposer: Tuning-Free Multi-Subject Image Generation with Localized Attention
Xiao, G.; Yin, T.; Freeman, W. T.; Durand, F.; and Han, S. 2023 · 2023
Closest in time.