Fetching the paper…
Reading the bibliography…
Tuning-free approaches adapting large-scale pre-trained video diffusion models for identity-preserving text-to-video generation (IPT2V) have gained popularity recently due to their efficacy and scalability.
“Two-frame motion estimation based on polynomial expansion”
Gunnar Farnebäck · 2003
Earlier work this paper cites.
“U-net: Convolutional networks for biomedical image segmentation”
Olaf Ronneberger, Philipp Fischer and Thomas Brox · 2015
Earlier work this paper cites.
“Gans trained by a two time-scale update rule converge to a local nash equilibrium”
Martin Heusel et al · 2017
Earlier work this paper cites.
“Learning a model of facial shape and expression from 4D scans.”
Tianye Li et al · 2017
Earlier work this paper cites.
“East: an efficient and accurate scene text detector”
Xinyu Zhou et al · 2017
Earlier work this paper cites.
“X2face: A network for controlling face generation using images, audio, and pose codes”
Olivia Wiles, A Koepke and Andrew Zisserman · 2018
Earlier work this paper cites.
“Arcface: Additive angular margin loss for deep face recognition”
Jiankang Deng, Jia Guo, Niannan Xue and Stefanos Zafeiriou · 2019
Earlier work this paper cites.
“Spiralnet++: A fast and highly efficient mesh convolution operator”
Shunwang Gong, Lei Chen, Michael Bronstein and Stefanos Zafeiriou · 2019
Earlier work this paper cites.
“Mediapipe: A framework for building perception pipelines”
Camillo Lugaresi et al · 2019
Earlier work this paper cites.
“RetinaFace: Single-Shot Multi-Level Face Localisation in the Wild”
Jiankang Deng et al · 2020
Earlier work this paper cites.
“Retinaface: Single-shot multi-level face localisation in the wild”
Jiankang Deng et al · 2020
Earlier work this paper cites.
“Mesh guided one-shot face reenactment using graph convolutional networks”
Guangming Yao, Yi Yuan, Tianjia Shao and Kun Zhou · 2020
Earlier work this paper cites.
“Learning an Animatable Detailed 3D Face Model from In-The-Wild Images”
Yao Feng, Haiwen Feng, Michael. Black and Timo Bolkart · 2021
Earlier work this paper cites.
“Learning an animatable detailed 3D face model from in-the-wild images”
Yao Feng, Haiwen Feng, Michael Black and Timo Bolkart · 2021
Earlier work this paper cites.
“Lora: Low-rank adaptation of large language models”
Edward Hu et al · 2021
Earlier work this paper cites.
“Learning transferable visual models from natural language supervision”
Alec Radford et al · 2021
Earlier work this paper cites.
“One-shot free-view neural talking-head synthesis for video conferencing”
Ting-Chun Wang, Arun Mallya and Ming-Yu Liu · 2021
Cited alongside, same era.
“An image is worth one word: Personalizing text-to-image generation using textual inversion”
Rinon Gal et al · 2022
Cited alongside, same era.
“Realistic one-shot mesh-based head avatars”
Taras Khakhulin, Vanessa Sklyarova, Victor Lempitsky and Egor Zakharov · 2022
Cited alongside, same era.
“Real-time scene text detection with differentiable binarization and adaptive scale fusion”
Minghui Liao et al · 2022
Cited alongside, same era.
“CelebV-HQ: A Large-Scale Video Facial Attributes Dataset”
Hao Zhu et al · 2022
Cited alongside, same era.
“A morphable model for the synthesis of 3D faces”
Volker Blanz and Thomas Vetter · 2023
“Echomimic: Lifelike audio-driven portrait animations through editable landmark conditions”
Zhiyuan Chen et al · 2024
Later among the works it cites.
“MagicMirror: Fast and High-Quality Avatar Generation with a Constrained Search Space”
Armand Comas-Massagué et al · 2024
Later among the works it cites.
“Pulid: Pure and lightning id customization via contrastive alignment”
Zinan Guo et al · 2024
Later among the works it cites.
Junjie He, Yifeng Geng and Liefeng Bo · 2024
Later among the works it cites.
“Id-animator: Zero-shot identity-preserving human video generation”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Stable video diffusion: Scaling latent video diffusion models to large datasets”
Andreas Blattmann et al · 2023
Cited alongside, same era.
“Animatediff: Animate your personalized text-to-image diffusion models without specific tuning”
Yuwei Guo et al · 2023
Cited alongside, same era.
“Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models”
Junnan Li, Dongxu Li, Silvio Savarese and Steven Hoi · 2023
Cited alongside, same era.
“Dpe: Disentanglement of pose and expression for general video portrait editing”
Youxin Pang et al · 2023
Cited alongside, same era.
“Scalable diffusion models with transformers”
William Peebles and Saining Xie · 2023
Cited alongside, same era.
“Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation”
Nataniel Ruiz et al · 2023
Cited alongside, same era.
Xuanhua He et al · 2024
Later among the works it cites.
“HunyuanVideo: A Systematic Framework For Large Video Generative Models”
Weijie Kong et al · 2024
Later among the works it cites.
“CVTHead: One-shot Controllable Head Avatar with Vertex-feature Transformer”
Haoyu Ma et al · 2024
Later among the works it cites.
“Magic-me: Identity-specific video customized diffusion”
Ze Ma et al · 2024
Later among the works it cites.
“OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation”
Kepan Nan et al · 2024
Later among the works it cites.
“Movie gen: A cast of media foundation models”
Adam Polyak et al · 2024
Later among the works it cites.
“Instantid: Zero-shot identity-preserving generation in seconds”
Qixun Wang et al · 2024
Later among the works it cites.
Tao Wu et al · 2024
Later among the works it cites.
“Cogvideox: Text-to-video diffusion models with an expert transformer”
Zhuoyi Yang et al · 2024
Later among the works it cites.
“Representation alignment for generation: Training diffusion transformers is easier than you think”
Sihyun Yu et al · 2024
Later among the works it cites.
“Identity-Preserving Text-to-Video Generation by Frequency Decomposition”
Shenghai Yuan et al · 2024
Later among the works it cites.
“Open-Sora: Democratizing Efficient Video Production for All”, 2024
Zangwei Zheng et al · 2024
Later among the works it cites.