Fetching the paper…
Reading the bibliography…
Vivid talking face generation holds immense potential applications across diverse multimedia domains, such as film and game production.
DiffTalk: Crafting Diffusion Models for Generalized Audio-Driven Portraits Animation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 1982–1991
Shuai Shen, Wenliang Zhao, Zibin Meng, Wanhua Li, Zheng Zhu, Jie Zhou, and Jiwen Lu. 2023 · 1991
Earlier work this paper cites.
Out of time: automated lip sync in the wild. In Workshop on Multi-view Lip-reading, ACCV. Springer, 251–263
Joon Son Chung and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017 · 2017
Earlier work this paper cites.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen. 2017 · 2017
Earlier work this paper cites.
A deep learning approach for generalized speech animation
Sarah Taylor, Taehwan Kim, Yisong Yue, Moshe Mahler, James Krahe, Anastasio Garcia Rodriguez, Jessica Hodgins, and Iain Matthews. 2017 · 2017
Earlier work this paper cites.
Lip movements generation at a glance. In Proceedings of the European conference on computer vision (ECCV) . 520–535
Lele Chen, Zhiheng Li, Ross K Maddox, Zhiyao Duan, and Chenliang Xu. 2018 · 2018
Earlier work this paper cites.
Hierarchical cross-modal talking face generation with dynamic pixel-wise loss. In CVPR . 7832–7841
Lele Chen, Ross K Maddox, Zhiyao Duan, and Chenliang Xu. 2019 · 2019
Earlier work this paper cites.
Mediapipe: A framework for building perception pipelines
Camillo Lugaresi, Jiuqiang Tang, Hadon Nash, Chris McClanahan, Esha Uboweja, Michael Hays, Fan Zhang, Chuo-Ling Chang, Ming Guang Yong, Juhyun Lee, et al · 2019
Earlier work this paper cites.
End-to-end speech-driven facial animation with temporal GANs. In CVPR Workshop
Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic. 2019 · 2019
Earlier work this paper cites.
Talking face generation by adversarially disentangled audio-visual representation. In AAAI , Vol. 33. 9299–9306
Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang. 2019 · 2019
Earlier work this paper cites.
Neural head reenactment with latent pose descriptors. In CVPR . 13786–13795
Egor Burkov, Igor Pasechnik, Artur Grigorev, and Victor Lempitsky. 2020 · 2020
Earlier work this paper cites.
Talking-head generation with rhythmic head motion. In ECCV . Springer, 35–51
Lele Chen, Guofeng Cui, Celong Liu, Zhong Li, Ziyi Kou, Yi Xu, and Chenliang Xu. 2020 · 2020
Earlier work this paper cites.
Flnet: Landmark driven fetching and learning network for faithful talking facial animation synthesis. In AAAI , Vol. 34. 10861–10868
Kuangxiao Gu, Yuqian Zhou, and Thomas Huang. 2020 · 2020
Cited alongside, same era.
Towards fast, accurate and stable 3d dense face alignment. In ECCV . Springer, 152–168
Jianzhu Guo, Xiangyu Zhu, Yang Yang, Fan Yang, Zhen Lei, and Stan Z Li. 2020 · 2020
Cited alongside, same era.
A lip sync expert is all you need for speech to lip generation in the wild. In ACM MM . 484–492
KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar. 2020 · 2020
Cited alongside, same era.
Mead: A large-scale audio-visual dataset for emotional talking-face generation. In ECCV . Springer, 700–717
Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, and Chen Change Loy. 2020 · 2020
Cited alongside, same era.
One-shot identity-preserving portrait reenactment
Sitao Xiang, Yuming Gu, Pengda Xiang, Mingming He, Koki Nagano, Haiwei Chen, and Hao Li. 2020 · 2020
Faceformer: Speech-driven 3d facial animation with transformers. In CVPR . 18770–18780
Yingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang, and Taku Komura. 2022 · 2022
Later among the works it cites.
Depth-aware generative adversarial network for talking head video generation. In CVPR . 3397–3406
Fa-Ting Hong, Longhao Zhang, Li Shen, and Dan Xu. 2022 · 2022
Later among the works it cites.
Eamm: One-shot emotional talking face via audio-based emotion-aware motion model. In ACM SIGGRAPH . 1–10
Xinya Ji, Hang Zhou, Kaisiyuan Wang, Qianyi Wu, Wayne Wu, Feng Xu, and Xun Cao. 2022 · 2022
Later among the works it cites.
Semantic-aware implicit neural audio-driven video portrait generation. In European Conference on Computer Vision . Springer, 106–125
Xian Liu, Yinghao Xu, Qianyi Wu, Hang Zhou, Wayne Wu, and Bolei Zhou. 2022 · 2022
Later among the works it cites.
Everybody’s talkin’: Let me talk as you want
Linsen Song, Wayne Wu, Chen Qian, Ran He, and Chen Change Loy. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Makelttalk: speaker-aware talking-head animation
Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li. 2020 · 2020
Cited alongside, same era.
Audio-driven emotional video portraits. In CVPR . 14080–14089
Xinya Ji, Hang Zhou, Kaisiyuan Wang, Wayne Wu, Chen Change Loy, Xun Cao, and Feng Xu. 2021 · 2021
Cited alongside, same era.
Li-net: Large-pose identity-preserving face reenactment network. In ICME . IEEE, 1–6
Jin Liu, Peng Chen, Tao Liang, Zhaoxing Li, Cai Yu, Shuqiao Zou, Jiao Dai, and Jizhong Han. 2021 · 2021
Cited alongside, same era.
Motion representations for articulated animation. In CVPR . 13653–13662
Aliaksandr Siarohin, Oliver J Woodford, Jian Ren, Menglei Chai, and Sergey Tulyakov. 2021 · 2021
Cited alongside, same era.
Audio2head: Audio-driven one-shot talking-head generation with natural head motion
Suzhen Wang, Lincheng Li, Yu Ding, Changjie Fan, and Xin Yu. 2021a · 2021
Cited alongside, same era.
Pose-controllable talking face generation by implicitly modularized audio-visual representation. In CVPR . 4176–4186
Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu. 2021 · 2021
Cited alongside, same era.
Megaportraits: One-shot megapixel neural head avatars. In ACM MM . 2663–2671
Nikita Drobyshev, Jenya Chelishev, Taras Khakhulin, Aleksei Ivakhnenko, Victor Lempitsky, and Egor Zakharov. 2022 · 2022
Cited alongside, same era.
Latent image animator: Learning to animate images via latent space navigation
Yaohui Wang, Di Yang, Francois Bremond, and Antitza Dantcheva. 2022 · 2022
Later among the works it cites.
Dae-talker: High fidelity speech-driven talking face generation with diffusion autoencoder. In Proceedings of the 31st ACM International Conference on Multimedia . 4281–4289
Chenpeng Du, Qi Chen, Tianyu He, Xu Tan, Xie Chen, Kai Yu, Sheng Zhao, and Jiang Bian. 2023 · 2023
Later among the works it cites.
Space: Speech-driven portrait animation with controllable expression. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 20914–20923
Siddharth Gururani, Arun Mallya, Ting-Chun Wang, Rafael Valle, and Ming-Yu Liu. 2023 · 2023
Later among the works it cites.
Emmn: Emotional motion memory network for audio-driven emotional talking face generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 22146–22156
Shuai Tan, Bin Ji, and Ye Pan. 2023 · 2023
Later among the works it cites.
Speech-Driven 3D Face Animation with Composite and Regional Facial Movements. In Proceedings of the 31st ACM International Conference on Multimedia . 6822–6830
Haozhe Wu, Songtao Zhou, Jia Jia, Junliang Xing, Qi Wen, and Xiang Wen. 2023 · 2023
Later among the works it cites.
High-fidelity Generalized Emotional Talking Face Generation with Multi-modal Emotion Space Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 6609–6619
Chao Xu, Junwei Zhu, Jiangning Zhang, Yue Han, Wenqing Chu, Ying Tai, Chengjie Wang, Zhifeng Xie, and Yong Liu. 2023 · 2023
Later among the works it cites.