Fetching the paper…
Reading the bibliography…
Customized video generation aims to generate high-quality videos guided by text prompts and subject's reference images.
Multi-concept customization of text-to-image diffusion
Kumari, N.; Zhang, B.; Zhang, R.; Shechtman, E.; and Zhu, J.-Y. 2023 · 1941
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J.; Meng, C.; and Ermon, S. 2020 · 2010
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J.; Weiss, E.; Maheswaranathan, N.; and Ganguli, S. 2015 · 2015
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Schuhmann, C.; Vencu, R.; Beaumont, R.; Kaczmarczyk, R.; Mullis, C.; Katta, A.; Coombes, T.; Jitsev, J.; and Komatsuzaki, A. 2021 · 2021
Earlier work this paper cites.
Latent Video Diffusion Models for High-Fidelity Video Generation with Arbitrary Lengths
He, Y.; Yang, T.; Zhang, Y.; Shan, Y.; and Chen, Q. 2022 · 2022
Earlier work this paper cites.
Video diffusion models
Ho, J.; Salimans, T.; Gritsenko, A.; Chan, W.; Norouzi, M.; and Fleet, D. J. 2022 · 2022
Earlier work this paper cites.
Hierarchical text-conditional image generation with clip latents
Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022 · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Earlier work this paper cites.
Dino: Detr with improved denoising anchor boxes for end-to-end object detection
Zhang, H.; Li, F.; Liu, S.; Zhang, L.; Su, H.; Zhu, J.; Ni, L. M.; and Shum, H.-Y. 2022 · 2022
Earlier work this paper cites.
Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing
Cao, M.; Wang, X.; Qi, Z.; Shan, Y.; Qie, X.; and Zheng, Y. 2023 · 2023
Earlier work this paper cites.
VideoCrafter1: Open Diffusion Models for High-Quality Video Generation
Chen, H.; Xia, M.; He, Y.; Zhang, Y.; Cun, X.; Yang, S.; Xing, J.; Liu, Y.; Chen, Q.; Wang, X.; et al. 2023 · 2023
Earlier work this paper cites.
Structure and content-guided video synthesis with diffusion models
Esser, P.; Chiu, J.; Atighehchian, P.; Granskog, J.; and Germanidis, A. 2023 · 2023
Cited alongside, same era.
An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion
Gal, R.; Alaluf, Y.; Atzmon, Y.; Patashnik, O.; Bermano, A. H.; Chechik, G.; and Cohen-or, D. 2023 · 2023
Cited alongside, same era.
Preserve your own correlation: A noise prior for video diffusion models
Ge, S.; Nah, S.; Liu, G.; Poon, T.; Tao, A.; Catanzaro, B.; Jacobs, D.; Huang, J.-B.; Liu, M.-Y.; and Balaji, Y. 2023 · 2023
Cited alongside, same era.
Svdiff: Compact parameter space for diffusion fine-tuning
Han, L.; Li, Y.; Zhang, H.; Milanfar, P.; Metaxas, D.; and Yang, F. 2023 · 2023
Cited alongside, same era.
Animate-a-story: Storytelling with retrieval-augmented video generation
He, Y.; Xia, M.; Chen, H.; Cun, X.; Gong, Y.; Xing, J.; Zhang, Y.; Wang, X.; Weng, C.; Shan, Y.; et al. 2023 · 2023
Cited alongside, same era.
GVDIFF: Grounded Text-to-Video Generation with Diffusion Models
Dou, H.; Li, R.; Su, W.; and Li, X. 2024 · 2024
Closest in time.
AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
Guo, Y.; Yang, C.; Rao, A.; Liang, Z.; Wang, Y.; Qiao, Y.; Agrawala, M.; Lin, D.; and Dai, B. 2024 · 2024
Closest in time.
ID-Animator: Zero-Shot Identity-Preserving Human Video Generation
He, X.; Liu, Q.; Qian, S.; Wang, X.; Hu, T.; Cao, K.; Yan, K.; Zhou, M.; and Zhang, J. 2024 · 2024
Closest in time.
A Survey of Multimodal Controllable Diffusion Models
Jiang, R.; Zheng, G.-C.; Li, T.; Yang, T.-R.; Wang, J.-D.; and Li, X. 2024 · 2024
Closest in time.
Blip-diffusion: Pre-trained subject representation for controllable text-to-image generation and editing
Li, D.; Li, J.; and Hoi, S. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
VideoBooth: Diffusion-based Video Generation with Image Prompts
Jiang, Y.; Wu, T.; Yang, S.; Si, C.; Lin, D.; Qiao, Y.; Loy, C. C.; and Liu, Z. 2023 · 2023
Cited alongside, same era.
Self-Paced Multi-Grained Cross-Modal Interaction Modeling for Referring Expression Comprehension
Miao, P.; Su, W.; Wang, G.; Li, X.; and Li, X. 2023 · 2023
Cited alongside, same era.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Ruiz, N.; Li, Y.; Jampani, V.; Pritch, Y.; Rubinstein, M.; and Aberman, K. 2023 · 2023
Cited alongside, same era.
Elite: Encoding visual concepts into textual embeddings for customized text-to-image generation
Wei, Y.; Zhang, Y.; Ji, Z.; Bai, J.; Zhang, L.; and Zuo, W. 2023 · 2023
Cited alongside, same era.
Bridging cross-task protocol inconsistency for distillation in dense object detection
Yang, L.; Zhou, X.; Li, X.; Qiao, L.; Li, Z.; Yang, Z.; Wang, G.; and Li, X. 2023 · 2023
Cited alongside, same era.
Video probabilistic diffusion models in projected latent space
Yu, S.; Sohn, K.; Kim, S.; and Shin, J. 2023 · 2023
Cited alongside, same era.
Videoassembler: Identity-consistent video generation with reference entities using diffusion model
Zhao, H.; Lu, T.; Gu, J.; Zhang, X.; Wu, Z.; Xu, H.; and Jiang, Y.-G. 2023 · 2023
Cited alongside, same era.
Towards Understanding Cross and Self-Attention in Stable Diffusion for Text-Guided Image Editing
Liu, B.; Wang, C.; Cao, T.; Jia, K.; and Huang, J. 2024 · 2024
Closest in time.
Subject-diffusion: Open domain personalized text-to-image generation without test-time fine-tuning
Ma, J.; Liang, J.; Chen, C.; and Lu, H. 2024 · 2024
Closest in time.
Dreammatcher: Appearance matching self-attention for semantically-consistent text-to-image personalization
Nam, J.; Kim, H.; Lee, D.; Jin, S.; Kim, S.; and Chang, S. 2024 · 2024
Closest in time.
Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models
Ruiz, N.; Li, Y.; Jampani, V.; Wei, W.; Hou, T.; Pritch, Y.; Wadhwa, N.; Rubinstein, M.; and Aberman, K. 2024 · 2024
Closest in time.
Instantbooth: Personalized text-to-image generation without test-time finetuning
Shi, J.; Xiong, W.; Lin, Z.; and Jung, H. J. 2024 · 2024
Closest in time.
Dreamvideo: Composing your dream videos with customized subject and motion
Wei, Y.; Zhang, S.; Qing, Z.; Yuan, H.; Liu, Z.; Liu, Y.; Zhang, Y.; Zhou, J.; and Shan, H. 2024 · 2024
Closest in time.
Fastcomposer: Tuning-free multi-subject image generation with localized attention
Xiao, G.; Yin, T.; Freeman, W. T.; Durand, F.; and Han, S. 2024 · 2024
Closest in time.
Inserting Anybody in Diffusion Models via Celeb Basis
Yuan, G.; Cun, X.; Zhang, Y.; Li, M.; Qi, C.; Wang, X.; Shan, Y.; and Zheng, H. 2024 · 2024
Closest in time.
Real-world image variation by aligning diffusion inversion chain
Zhang, Y.; Xing, J.; Lo, E.; and Jia, J. 2024 · 2024
Closest in time.