Fetching the paper…
Reading the bibliography…
The objective of face animation is to generate dynamic and expressive talking head videos from a single reference face, utilizing driving conditions derived from either video or audio inputs.
Difftalk: Crafting diffusion models for generalized audio-driven portraits animation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 1982–1991
Shuai Shen, Wenliang Zhao, Zibin Meng, Wanhua Li, Zheng Zhu, Jie Zhou, and Jiwen Lu. 2023 · 1991
Earlier work this paper cites.
A lip sync expert is all you need for speech to lip generation in the wild. In Proceedings of the 28th ACM international conference on multimedia . 484–492
KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar. 2020 · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision. In International conference on machine learning . PMLR, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 3661–3670
Zhimeng Zhang, Lincheng Li, Yu Ding, and Changjie Fan. 2021 · 2021
Earlier work this paper cites.
Megaportraits: One-shot megapixel neural head avatars. In Proceedings of the 30th ACM International Conference on Multimedia . 2663–2671
Nikita Drobyshev, Jenya Chelishev, Taras Khakhulin, Aleksei Ivakhnenko, Victor Lempitsky, and Egor Zakharov. 2022 · 2022
Earlier work this paper cites.
Wenxuan Zhang, Xiaodong Cun, Xuan Wang, Yong Zhang, Xi Shen, Yu Guo, Ying Shan, and Fei Wang. 2022 · 2022
Earlier work this paper cites.
CelebV-HQ: A Large-Scale Video Facial Attributes Dataset. In ECCV
Hao Zhu, Wayne Wu, Wentao Zhu, Liming Jiang, Siwei Tang, Li Zhang, Ziwei Liu, and Chen Change Loy. 2022 · 2022
Cited alongside, same era.
DaGAN++: Depth-Aware Generative Adversarial Network for Talking Head Video Generation
Fa-Ting Hong, , Li Shen, and Dan Xu. 2023 · 2023
Cited alongside, same era.
Implicit Identity Representation Conditioned Memory Compensation Network for Talking Head video Generation. In ICCV
Fa-Ting Hong and Dan Xu. 2023 · 2023
Cited alongside, same era.
Animate anyone: Consistent and controllable image-to-video synthesis for character animation
Li Hu, Xin Gao, Peng Zhang, Ke Sun, Bang Zhang, and Liefeng Bo. 2023 · 2023
Cited alongside, same era.
Cross-identity video motion retargeting with joint transformation and synthesis. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision . 412–422
Haomiao Ni, Yihao Liu, Sharon X Huang, and Yuan Xue. 2023 · 2023
Effective whole-body pose estimation with two-stages distillation. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 4210–4220
Zhendong Yang, Ailing Zeng, Chun Yuan, and Yu Li. 2023 · 2023
Later among the works it cites.
GeneFace++: Generalized and Stable Real-Time Audio-Driven 3D Talking Face Generation
Zhenhui Ye, Jinzheng He, Ziyue Jiang, Rongjie Huang, Jiawei Huang, Jinglin Liu, Yi Ren, Xiang Yin, Zejun Ma, and Zhou Zhao. 2023 · 2023
Later among the works it cites.
Linrui Tian, Qi Wang, Bang Zhang, and Liefeng Bo. 2024 · 2024
Closest in time.
Aniportrait: Audio-driven synthesis of photorealistic portrait animation
Huawei Wei, Zejun Yang, and Zhisheng Wang. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Zhenhui Ye, Tianyun Zhong, Yi Ren, Jiaqi Yang, Weichuang Li, Jiawei Huang, Ziyue Jiang, Jinzheng He, Rongjie Huang, Jinglin Liu, et al · 2024
Closest in time.