Fetching the paper…
Reading the bibliography…
Recently, the success of text-to-image synthesis has greatly advanced the development of identity customization techniques, whose main goal is to produce realistic identity-specific photographs based on text prompts and reference face images.
Mediapipe: A framework for building perception pipelines
Lugaresi, C.; Tang, J.; Nash, H.; McClanahan, C.; Uboweja, E.; Hays, M.; Zhang, F.; Chang, C.-L.; Yong, M. G.; Lee, J.; et al. 2019 · 1906
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J.; Meng, C.; and Ermon, S. 2020 · 2010
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
Schroff, F.; Kalenichenko, D.; and Philbin, J. 2015 · 2015
Earlier work this paper cites.
Joint face detection and alignment using multitask cascaded convolutional networks
Zhang, K.; Zhang, Z.; Li, Z.; and Qiao, Y. 2016 · 2016
Earlier work this paper cites.
Progressive growing of gans for improved quality, stability, and variation
Karras, T.; Aila, T.; Laine, S.; and Lehtinen, J. 2017 · 2017
Earlier work this paper cites.
Sod-mtgan: Small object detection via multi-task generative adversarial network
Bai, Y.; Zhang, Y.; Ding, M.; and Ghanem, B. 2018 · 2018
Earlier work this paper cites.
Bisenet: Bilateral segmentation network for real-time semantic segmentation
Yu, C.; Wang, J.; Peng, C.; Gao, C.; Yu, G.; and Sang, N. 2018 · 2018
Earlier work this paper cites.
Better to follow, follow to be better: Towards precise supervision of feature super-resolution for small object detection
Noh, J.; Bae, W.; Lee, W.; Seo, J.; and Kim, G. 2019 · 2019
Earlier work this paper cites.
Learning an animatable detailed 3D face model from in-the-wild images
Feng, Y.; Feng, H.; Black, M. J.; and Bolkart, T. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Earlier work this paper cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Nichol, A.; Dhariwal, P.; Ramesh, A.; Shyam, P.; Mishkin, P.; McGrew, B.; Sutskever, I.; and Chen, M. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
An image is worth one word: Personalizing text-to-image generation using textual inversion
Gal, R.; Alaluf, Y.; Atzmon, Y.; Patashnik, O.; Bermano, A. H.; Chechik, G.; and Cohen-Or, D. 2022 · 2022
Cited alongside, same era.
Classifier-free diffusion guidance
Ho, J.; and Salimans, T. 2022 · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022 · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models
Ye, H.; Zhang, J.; Liu, S.; Han, X.; and Yang, W. 2023 · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
Zhang, L.; Rao, A.; and Agrawala, M. 2023 · 2023
Later among the works it cites.
IDAdapter: Learning Mixed Features for Tuning-Free Personalization of Text-to-Image Models
Cui, S.; Guo, J.; An, X.; Deng, J.; Zhao, Y.; Wei, X.; and Feng, Z. 2024 · 2024
Closest in time.
PuLID: Pure and Lightning ID Customization via Contrastive Alignment
Guo, Z.; Wu, Y.; Chen, Z.; Chen, L.; and He, Q. 2024 · 2024
Closest in time.
Caphuman: Capture your moments in parallel universes
Liang, C.; Ma, F.; Zhu, L.; Deng, Y.; and Yang, Y. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E. L.; Ghasemipour, K.; Gontijo Lopes, R.; Karagol Ayan, B.; Salimans, T.; et al. 2022 · 2022
Cited alongside, same era.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C.; Beaumont, R.; Vencu, R.; Gordon, C.; Wightman, R.; Cherti, M.; Coombes, T.; Katta, A.; Mullis, C.; Wortsman, M.; et al. 2022 · 2022
Cited alongside, same era.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Podell, D.; English, Z.; Lacey, K.; Blattmann, A.; Dockhorn, T.; Müller, J.; Penna, J.; and Rombach, R. 2023 · 2023
Cited alongside, same era.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Ruiz, N.; Li, Y.; Jampani, V.; Pritch, Y.; Rubinstein, M.; and Aberman, K. 2023 · 2023
Cited alongside, same era.
Fastcomposer: Tuning-free multi-subject image generation with localized attention
Xiao, G.; Yin, T.; Freeman, W. T.; Durand, F.; and Han, S. 2023 · 2023
Cited alongside, same era.
Paint by example: Exemplar-based image editing with diffusion models
Yang, B.; Gu, S.; Zhang, B.; Zhang, T.; Chen, X.; Sun, X.; Chen, D.; and Wen, F. 2023 · 2023
Cited alongside, same era.
InstantFamily: Masked Attention for Zero-shot Multi-ID Image Generation
Kim, C.; Lee, J.; Joung, S.; Kim, B.; and Baek, Y.-M. 2024a
Cited in the paper.
Ma, Y.; Liu, H.; Wang, H.; Pan, H.; He, Y.; Yuan, J.; Zeng, A.; Cai, C.; Shum, H.-Y.; Liu, W.; et al. 2024 · 2024
Closest in time.
T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models
Mou, C.; Wang, X.; Xie, L.; Wu, Y.; Zhang, J.; Qi, Z.; and Shan, Y. 2024 · 2024
Closest in time.
Portraitbooth: A versatile portrait model for fast identity-preserved personalization
Peng, X.; Zhu, J.; Jiang, B.; Tai, Y.; Luo, D.; Zhang, J.; Lin, W.; Jin, T.; Wang, C.; and Ji, R. 2024 · 2024
Closest in time.
Aniportrait: Audio-driven synthesis of photorealistic portrait animation
Wei, H.; Yang, Z.; and Wang, Z. 2024 · 2024
Closest in time.
MM-Diff: High-Fidelity Image Personalization via Multi-Modal Condition Integration
Wei, Z.; Su, Q.; Qin, L.; and Wang, W. 2024 · 2024
Closest in time.
FlashFace: Human Image Personalization with High-fidelity Identity Preservation
Zhang, S.; Huang, L.; Chen, X.; Zhang, Y.; Wu, Z.-F.; Feng, Y.; Wang, W.; Shen, Y.; Liu, Y.; and Luo, P. 2024 · 2024
Closest in time.