Fetching the paper…
Reading the bibliography…
Customized video generation aims to produce videos featuring specific subjects under flexible user-defined conditions, yet existing methods often struggle with identity consistency and limited input modalities.
General data protection regulation
P. Regulation · 2018
Earlier work this paper cites.
Arcface: Additive angular margin loss for deep face recognition
J. Deng, J. Guo, N. Xue, and S. Zafeiriou · 2019
Earlier work this paper cites.
Pyscenedetect
B. Castellano · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
An image is worth one word: Personalizing text-to-image generation using textual inversion
R. Gal, Y. Alaluf, Y. Atzmon, O. Patashnik, A. H. Bermano, G. Chechik, and D. Cohen-Or · 2022
Earlier work this paper cites.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
W. Hong, M. Ding, W. Zheng, X. Liu, and J. Tang · 2022
Earlier work this paper cites.
Flow matching for generative modeling
Y. Lipman, R. T. Chen, H. Ben-Hamu, M. Nickel, and M. Le · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Earlier work this paper cites.
J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, et al · 2023
Earlier work this paper cites.
High-fidelity and freely controllable talking head video generation
Y. Gao, Y. Zhou, J. Wang, X. Li, X. Ming, and Y. Lu · 2023
Earlier work this paper cites.
Dinov2: Learning robust visual features without supervision
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al · 2023
Earlier work this paper cites.
Robust speech recognition via large-scale weak supervision
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever · 2023
Earlier work this paper cites.
Facial geometric detail recovery via implicit representation
X. Ren, A. Lattas, B. Gecer, J. Deng, C. Ma, and X. Yang · 2023
Earlier work this paper cites.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
N. Ruiz, Y. Li, V. Jampani, Y. Pritch, M. Rubinstein, and K. Aberman · 2023
Earlier work this paper cites.
Still-moving: Customized video generation without customized video data
H. Chefer, S. Zada, R. Paiss, A. Ephrat, O. Tov, M. Rubinstein, L. Wolf, T. Dekel, T. Michaeli, and I. Mosseri · 2024
Cited alongside, same era.
Disenstudio: Customized multi-subject text-to-video generation with disentangled spatial control
H. Chen, X. Wang, Y. Zhang, Y. Zhou, Z. Zhang, S. Tang, and W. Zhu · 2024
Cited alongside, same era.
Scaling rectified flow transformers for high-resolution image synthesis
P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y. Levi, D. Lorenz, A. Sauer, F. Boesel, et al · 2024
Cited alongside, same era.
Id-animator: Zero-shot identity-preserving human video generation
X. He, Q. Liu, S. Qian, X. Wang, T. Hu, K. Cao, K. Yan, and J. Zhang · 2024
Cited alongside, same era.
Motionmaster: Training-free camera motion transfer for video generation
T. Hu, J. Zhang, R. Yi, Y. Wang, H. Huang, J. Weng, Y. Wang, and L. Ma · 2024
Florence-2: Advancing a unified representation for a variety of vision tasks
B. Xiao, H. Wu, W. Xu, X. Dai, H. Hu, Y. Lu, M. Zeng, C. Liu, and L. Yuan · 2024
Later among the works it cites.
Cogvideox: Text-to-video diffusion models with an expert transformer
Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y. Yang, W. Hong, X. Zhang, G. Feng, et al · 2024
Later among the works it cites.
Identity-preserving text-to-video generation by frequency decomposition
S. Yuan, J. Huang, X. He, Y. Ge, Y. Shi, L. Chen, J. Luo, and L. Yuan · 2024
Later among the works it cites.
Multi-subject open-set personalization in video generation
T.-S. Chen, A. Siarohin, W. Menapace, Y. Fang, K. S. Lee, I. Skorokhodov, K. Aberman, J.-Y. Zhu, M.-H. Yang, and S. Tulyakov · 2025
Closest in time.
Skyreels-a2: Compose anything in video diffusion transformers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sonic: Shifting focus to global audio perception in portrait animation
X. Ji, X. Hu, Z. Xu, J. Zhu, C. Lin, Q. He, J. Zhang, D. Luo, Y. Chen, Q. Lin, et al · 2024
Cited alongside, same era.
Yolov11: An overview of the key architectural enhancements
R. Khanam and M. Hussain · 2024
Cited alongside, same era.
Hunyuanvideo: A systematic framework for large video generative models
W. Kong, Q. Tian, Z. Zhang, R. Min, Z. Dai, J. Zhou, J. Xiong, X. Li, B. Wu, J. Zhang, et al · 2024
Cited alongside, same era.
Sora: A review on background, technology, limitations, and opportunities of large vision models
Y. Liu, K. Zhang, Y. Li, Z. Yan, C. Gao, R. Chen, Z. Yuan, Y. Huang, H. Sun, J. Gao, et al · 2024
Cited alongside, same era.
Expressive talking avatars
Y. Pan, S. Tan, S. Cheng, Q. Lin, Z. Zeng, and K. Mitchell · 2024
Cited alongside, same era.
Sam 2: Segment anything in images and videos, 2024
N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V. Alwala, N. Carion, C.-Y. Wu, R. Girshick, P. Dollár, and C. Feichtenhofer · 2024
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu · 2024
Cited alongside, same era.
Z. Fei, D. Li, D. Qiu, J. Wang, Y. Dou, R. Wang, J. Xu, M. Fan, G. Chen, Y. Li, et al · 2025
Closest in time.
Y. Huang, Z. Yuan, Q. Liu, Q. Wang, X. Wang, R. Zhang, P. Wan, D. Zhang, and K. Gai · 2025
Closest in time.
Vace: All-in-one video creation and editing
Z. Jiang, Z. Han, C. Mao, J. Zhang, Y. Pan, and Y. Liu · 2025
Closest in time.
Mvportrait: Text-guided motion and emotion control for multi-view vivid portrait animation
Y. Lin, H. Fung, J. Xu, Z. Ren, A. S. Lau, G. Yin, and X. Li · 2025
Closest in time.
Phantom: Subject-consistent video generation via cross-modal alignment
L. Liu, T. Ma, B. Li, Z. Chen, J. Liu, Q. He, and X. Wu · 2025
Closest in time.
Movie gen: A cast of media foundation models, 2025
A. Polyak, A. Zohar, A. Brown, A. Tjandra, A. Sinha, A. Lee, A. Vyas, B. Shi, C.-Y. Ma, C.-Y. Chuang, D. Yan, D. Choudhary, D. Wang, G. Sethi, G. Pang, H. Ma, et al · 2025
Closest in time.
Wan: Open and advanced large-scale video generative models
A. Wang, B. Ai, B. Wen, C. Mao, C.-W. Xie, D. Chen, F. Yu, H. Zhao, J. Yang, J. Zeng, et al · 2025
Closest in time.
Customcrafter: Customized video generation with preserving motion and concept composition abilities
T. Wu, Y. Zhang, X. Wang, X. Zhou, G. Zheng, Z. Qi, Y. Shan, and X. Li · 2025
Closest in time.
Hunyuanportrait: Implicit condition control for enhanced portrait animation
Z. Xu, Z. Yu, Z. Zhou, J. Zhou, X. Jin, F.-T. Hong, X. Ji, J. Zhu, C. Cai, S. Tang, et al · 2025
Closest in time.