Fetching the paper…
Reading the bibliography…
For audio-driven visual dubbing, it remains a considerable challenge to uphold and highlight speaker's persona while synthesizing accurate lip synchronization.
A morphable model for the synthesis of 3D faces. In Proceedings of the 26th Annual Conference on Computer Graphics and Interactive Techniques (SIGGRAPH ’99) . ACM Press/Addison-Wesley Publishing Co., USA, 187–194
Volker Blanz and Thomas Vetter. 1999 · 1999
Earlier work this paper cites.
Visualizing data using t-SNE
Laurens Van der Maaten and Geoffrey Hinton. 2008 · 2008
Earlier work this paper cites.
Out of time: automated lip sync in the wild. In Workshop on Multi-view Lip-reading, ACCV
J. S. Chung and A. Zisserman. 2016 · 2016
Earlier work this paper cites.
Out of time: automated lip sync in the wild. In Computer Vision–ACCV 2016 Workshops: ACCV 2016 International Workshops, Taipei, Taiwan, November 20-24, 2016, Revised Selected Papers, Part II 13 . Springer, 251–263
Joon Son Chung and Andrew Zisserman. 2017 · 2016
Earlier work this paper cites.
Arbitrary style transfer in real-time with adaptive instance normalization. In Proceedings of the IEEE international conference on computer vision . 1501–1510
Xun Huang and Serge Belongie. 2017 · 2017
Earlier work this paper cites.
Synthesizing Obama: learning lip sync from audio
Supasorn Suwajanakorn, Steven M. Seitz, and Ira Kemelmacher-Shlizerman. 2017 · 2017
Earlier work this paper cites.
Attention is all you need. In NIPS
A Waswani, N Shazeer, N Parmar, J Uszkoreit, L Jones, A Gomez, L Kaiser, and I Polosukhin. 2017 · 2017
Earlier work this paper cites.
Face alignment in full pose range: A 3d total solution
Xiangyu Zhu, Xiaoming Liu, Zhen Lei, and Stan Z Li. 2017 · 2017
Earlier work this paper cites.
VoxCeleb2: Deep Speaker Recognition. In INTERSPEECH
J. S. Chung, A. Nagrani, and A. Zisserman. 2018 · 2018
Earlier work this paper cites.
The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In CVPR
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. 2018 · 2018
Earlier work this paper cites.
Capture, Learning, and Synthesis of 3D Speaking Styles
Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael Black. 2019 · 2019
Earlier work this paper cites.
Ganfit: Generative adversarial network fitting for high fidelity 3d face reconstruction. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 1155–1164
Baris Gecer, Stylianos Ploumpis, Irene Kotsia, and Stefanos Zafeiriou. 2019 · 2019
Earlier work this paper cites.
Towards automatic face-to-face translation. In Proceedings of the 27th ACM international conference on multimedia . 1428–1436
Prajwal KR, Rudrabha Mukhopadhyay, Jerin Philip, Abhishek Jha, Vinay Namboodiri, and CV Jawahar. 2019 · 2019
Earlier work this paper cites.
MediaPipe: A Framework for Perceiving and Processing Reality. In CVPR 2019
Camillo Lugaresi, Jiuqiang Tang, Hadon Nash, Chris McClanahan, Esha Uboweja, Michael Hays, Fan Zhang, Chuo-Ling Chang, Ming Yong, Juhyun Lee, Wan-Teh Chang, Wei Hua, Manfred Georg, and Matthias Grundmann. 2019 · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Earlier work this paper cites.
Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 8110–8119
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2020 · 2020
Earlier work this paper cites.
A Lip Sync Expert Is All You Need for Speech to Lip Generation In the Wild. In Proceedings of the 28th ACM International Conference on Multimedia (Seattle, WA, USA) (MM ’20) . Association for Computing Machinery, New York, NY, USA, 484–492
K R Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, and C.V. Jawahar. 2020 · 2020
Earlier work this paper cites.
Blindly assess image quality in the wild guided by a self-adaptive hyper network. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 3667–3676
Shaolin Su, Qingsen Yan, Yu Zhu, Cheng Zhang, Xin Ge, Jinqiu Sun, and Yanning Zhang. 2020 · 2020
Cited alongside, same era.
MakeltTalk: speaker-aware talking-head animation
Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li. 2020 · 2020
Cited alongside, same era.
Realistic talking face animation with speech-induced head motion. In Proceedings of the Twelfth Indian Conference on Computer Vision, Graphics and Image Processing . 1–9
Sandika Biswas, Sanjana Sinha, Dipanjan Das, and Brojeshwar Bhowmick. 2021 · 2021
Cited alongside, same era.
Headgan: One-shot neural head synthesis and editing. In Proceedings of the IEEE/CVF International conference on Computer Vision . 14398–14407
Michail Christos Doukas, Stefanos Zafeiriou, and Viktoriia Sharmanska. 2021 · 2021
Cited alongside, same era.
Learning an Animatable Detailed 3D Face Model from In-The-Wild Images
VFHQ: A High-Quality Dataset and Benchmark for Video Face Super-Resolution. In The IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
Liangbin Xie, Xintao Wang, Honglun Zhang, Chao Dong, and Ying Shan. 2022 · 2022
Later among the works it cites.
CelebV-HQ: A Large-Scale Video Facial Attributes Dataset. In ECCV
Hao Zhu, Wayne Wu, Wentao Zhu, Liming Jiang, Siwei Tang, Li Zhang, Ziwei Liu, and Chen Change Loy. 2022 · 2022
Later among the works it cites.
StyleSync: High-Fidelity Generalized and Personalized Lip Sync in Style-based Generator. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
Jiazhi Guan, Zhanwang Zhang, Hang Zhou, Tianshu HU, Kaisiyuan Wang, Dongliang He, Haocheng Feng, Jingtuo Liu, Errui Ding, Ziwei Liu, and Jingdong Wang. 2023 · 2023
Later among the works it cites.
Efficient region-aware neural radiance fields for high-fidelity talking portrait synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 7568–7578
Jiahe Li, Jiawei Zhang, Xiao Bai, Jun Zhou, and Lin Gu. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yao Feng, Haiwen Feng, Michael J. Black, and Timo Bolkart. 2021 · 2021
Cited alongside, same era.
Ad-nerf: Audio driven neural radiance fields for talking head synthesis. In Proceedings of the IEEE/CVF international conference on computer vision . 5784–5794
Yudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu, Hujun Bao, and Juyong Zhang. 2021 · 2021
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021 · 2021
Cited alongside, same era.
Live Speech Portraits: Real-Time Photorealistic Talking-Head Animation
Yuanxun Lu, Jinxiang Chai, and Xun Cao. 2021 · 2021
Cited alongside, same era.
Nerf: Representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. 2021 · 2021
Cited alongside, same era.
MeshTalk: 3D Face Animation From Speech Using Cross-Modality Disentanglement. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . 1173–1182
Alexander Richard, Michael Zollhöfer, Yandong Wen, Fernando de la Torre, and Yaser Sheikh. 2021 · 2021
Cited alongside, same era.
Imitating Arbitrary Talking Style for Realistic Audio-Driven Talking Face Synthesis. In Proceedings of the 29th ACM International Conference on Multimedia (Virtual Event, China) (MM ’21) . Association for Computing Machinery, New York, NY, USA, 1478–1486
Haozhe Wu, Jia Jia, Haoyu Wang, Yishun Dou, Chao Duan, and Qingshan Deng. 2021 · 2021
Cited alongside, same era.
Iterative Text-Based Editing of Talking-Heads Using Neural Retargeting
Xinwei Yao, Ohad Fried, Kayvon Fatahalian, and Maneesh Agrawala. 2021 · 2021
Cited alongside, same era.
Yifeng Ma, Shiwei Zhang, Jiayu Wang, Xiang Wang, Yingya Zhang, and Zhidong Deng. 2023b · 2023
Later among the works it cites.
DiffTalk: Crafting Diffusion Models for Generalized Audio-Driven Portraits Animation. In CVPR
Shuai Shen, Wenliang Zhao, Zibin Meng, Wanhua Li, Zheng Zhu, Jie Zhou, and Jiwen Lu. 2023 · 2023
Later among the works it cites.
Zhiyao Sun, Tian Lv, Sheng Ye, Matthieu Gaetan Lin, Jenny Sheng, Yu-Hui Wen, Minjing Yu, and Yong-jin Liu. 2023 · 2023
Later among the works it cites.
Imitator: Personalized Speech-driven 3D Facial Animation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . 20621–20631
Balamurugan Thambiraja, Ikhsanul Habibie, Sadegh Aliakbarian, Darren Cosker, Christian Theobalt, and Justus Thies. 2023 · 2023
Later among the works it cites.
Seeing What You Said: Talking Face Generation Guided by a Lip Reading Expert. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 14653–14662
Jiadong Wang, Xinyuan Qian, Malu Zhang, Robby T Tan, and Haizhou Li. 2023 · 2023
Later among the works it cites.
Speech-Driven 3D Face Animation with Composite and Regional Facial Movements. In Proceedings of the 31st ACM International Conference on Multimedia . 6822–6830
Haozhe Wu, Songtao Zhou, Jia Jia, Junliang Xing, Qi Wen, and Xiang Wen. 2023 · 2023
Later among the works it cites.
Codetalker: Speech-driven 3d facial animation with discrete motion prior. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 12780–12790
Jinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun, Jue Wang, and Tien-Tsin Wong. 2023 · 2023
Later among the works it cites.
GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis
Zhenhui Ye, Ziyue Jiang, Yi Ren, Jinglin Liu, Jinzheng He, and Zhou Zhao. 2023 · 2023
Later among the works it cites.
DINet: Deformation Inpainting Network for Realistic Face Visually Dubbing on High Resolution Video
Zhimeng Zhang, Zhipeng Hu, Wenjin Deng, Changjie Fan, Tangjie Lv, and Yu Ding. 2023 · 2023
Later among the works it cites.
Identity-preserving talking face generation with landmark and appearance priors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 9729–9738
Weizhi Zhong, Chaowei Fang, Yinqi Cai, Pengxu Wei, Gangming Zhao, Liang Lin, and Guanbin Li. 2023 · 2023
Later among the works it cites.
Make-Your-Anchor: A Diffusion-based 2D Avatar Generation Framework
Ziyao Huang, Fan Tang, Yong Zhang, Xiaodong Cun, Juan Cao, Jintao Li, and Tong-Yee Lee. 2024 · 2024
Closest in time.
SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
Ziqiao Peng, Wentao Hu, Yue Shi, Xiangyu Zhu, Xiaomei Zhang, Jun He, Hongyan Liu, and Zhaoxin Fan. 2024 · 2024
Closest in time.
Diffused heads: Diffusion models beat gans on talking-face generation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision . 5091–5100
Michał Stypułkowski, Konstantinos Vougioukas, Sen He, Maciej Zięba, Stavros Petridis, and Maja Pantic. 2024 · 2024
Closest in time.