Fetching the paper…
Reading the bibliography…
The paper introduces AniTalker, an innovative framework designed to generate lifelike talking faces from a single portrait.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. 2004 · 2004
Earlier work this paper cites.
Long short-term memory
Alex Graves and Alex Graves. 2012 · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling. 2013 · 2013
Earlier work this paper cites.
Deep Learning Face Attributes in the Wild. In Proceedings of International Conference on Computer Vision (ICCV)
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. 2015 · 2015
Earlier work this paper cites.
The importance of non-verbal communication
Deepika Phutela. 2015 · 2015
Earlier work this paper cites.
Out of time: automated lip sync in the wild. In Asian Conference on Computer Vision (ACCV) Workshops
Joon Son Chung and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2117–2125
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. 2017 · 2017
Earlier work this paper cites.
Voxceleb: a large-scale speaker identification dataset
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
Mutual Information Neural Estimation. In International Conference on Machine Learning (ICML) (Proceedings of Machine Learning Research, Vol. 80) , Jennifer Dy and Andreas Krause (Eds.). PMLR, 531–540
Mohamed Ishmael Belghazi, Aristide Baratin, Sai Rajeswar, Sherjil Ozair, Yoshua Bengio, Aaron Courville, and R Devon Hjelm. 2018 · 2018
Earlier work this paper cites.
Triplet loss in siamese network for object tracking. In Proceedings of the European conference on computer vision (ECCV) . 459–474
Xingping Dong and Jianbing Shen. 2018 · 2018
Earlier work this paper cites.
Additive margin softmax for face verification
Feng Wang, Jian Cheng, Weiyang Liu, and Haijun Liu. 2018 · 2018
Earlier work this paper cites.
X2face: A network for controlling face generation using images, audio, and pose codes. In Proceedings of the European conference on computer vision (ECCV) . 670–686
Olivia Wiles, A Koepke, and Andrew Zisserman. 2018 · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR)
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. 2018 · 2018
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. 2019 · 2019
Earlier work this paper cites.
Arcface: Additive angular margin loss for deep face recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 4690–4699
Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. 2019 · 2019
Earlier work this paper cites.
First order motion model for image animation
Aliaksandr Siarohin, Stéphane Lathuilière, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. 2019 · 2019
Earlier work this paper cites.
Club: A contrastive log-ratio upper bound of mutual information. In International Conference on Machine Learning (ICML) . PMLR, 1779–1788
Pengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu, Zhe Gan, and Lawrence Carin. 2020 · 2020
Earlier work this paper cites.
ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification
Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck. 2020 · 2020
Earlier work this paper cites.
Conformer: Convolution-augmented transformer for speech recognition
Anmol Gulati et al · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Earlier work this paper cites.
A lip sync expert is all you need for speech to lip generation in the wild. In Proceedings of the 28th ACM international conference on multimedia (ACM MM)
KR Prajwal et al · 2020
Cited alongside, same era.
Fastspeech 2: Fast and high-quality end-to-end text to speech
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2020 · 2020
Cited alongside, same era.
Denoising Diffusion Implicit Models. In International Conference on Learning Representations (ILCR)
Jiaming Song, Chenlin Meng, and Stefano Ermon. 2020 · 2020
Cited alongside, same era.
Learning an animatable detailed 3D face model from in-the-wild images
Yao Feng, Haiwen Feng, Michael J Black, and Timo Bolkart. 2021 · 2021
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021 · 2021
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Yuwei Guo, Ceyuan Yang, Anyi Rao, Yaohui Wang, Yu Qiao, Dahua Lin, and Bo Dai. 2023 · 2023
Later among the works it cites.
Animate anyone: Consistent and controllable image-to-video synthesis for character animation
Li Hu, Xin Gao, Peng Zhang, Ke Sun, Bang Zhang, and Liefeng Bo. 2023 · 2023
Later among the works it cites.
Dreamtalk: When expressive talking head generation meets diffusion probabilistic models
Yifeng Ma, Shiwei Zhang, Jiayu Wang, Xiang Wang, Yingya Zhang, and Zhidong Deng. 2023 · 2023
Later among the works it cites.
Dpe: Disentanglement of pose and expression for general video portrait editing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 427–436
Youxin Pang, Yong Zhang, Weize Quan, Yanbo Fan, Xiaodong Cun, Ying Shan, and Dong-ming Yan. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Pixel-in-pixel net: Towards efficient facial landmark detection in the wild
Haibo Jin, Shengcai Liao, and Ling Shao. 2021 · 2021
Cited alongside, same era.
Audio2head: Audio-driven one-shot talking-head generation with natural head motion
Suzhen Wang, Lincheng Li, Yu Ding, Changjie Fan, and Xin Yu. 2021a · 2021
Cited alongside, same era.
Superb: Speech processing universal performance benchmark
Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y Lin, Andy T Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, et al · 2021
Cited alongside, same era.
Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
Zhimeng Zhang, Lincheng Li, Yu Ding, and Changjie Fan. 2021 · 2021
Cited alongside, same era.
Pose-controllable talking face generation by implicitly modularized audio-visual representation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (CVPR)
Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu. 2021 · 2021
Cited alongside, same era.
Makelttalk: speaker-aware talking-head animation
Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li. 2020 · 2021
Cited alongside, same era.
Emoca: Emotion driven monocular face capture and animation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 20311–20322
Radek Daněček, Michael J Black, and Timo Bolkart. 2022 · 2022
Cited alongside, same era.
Inkyu Park and Jaewoong Cho. 2023 · 2023
Later among the works it cites.
Emotalk: Speech-driven emotional disentanglement for 3d face animation. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 20687–20697
Ziqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu, Xiangyu Zhu, Jun He, Hongyan Liu, and Zhaoxin Fan. 2023 · 2023
Later among the works it cites.
DiffTalk: Crafting Diffusion Models for Generalized Audio-Driven Portraits Animation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
Shuai Shen, Wenliang Zhao, Zibin Meng, Wanhua Li, Zheng Zhu, Jie Zhou, and Jiwen Lu. 2023 · 2023
Later among the works it cites.
Face Animation with an Attribute-Guided Diffusion Model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 628–637
Bohan Zeng, Xuhui Liu, Sicheng Gao, Boyu Liu, Hong Li, Jianzhuang Liu, and Baochang Zhang. 2023 · 2023
Later among the works it cites.
Chenxu Zhang, Chao Wang, Jianfeng Zhang, Hongyi Xu, Guoxian Song, You Xie, Linjie Luo, Yapeng Tian, Xiaohu Guo, and Jiashi Feng. 2023c · 2023
Later among the works it cites.
DINet: Deformation Inpainting Network for Realistic Face Visually Dubbing on High Resolution Video
Zhimeng Zhang et al · 2023
Later among the works it cites.
Identity-Preserving Talking Face Generation with Landmark and Appearance Priors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
Weizhi Zhong, Chaowei Fang, Yinqi Cai, Pengxu Wei, Gangming Zhao, Liang Lin, and Guanbin Li. 2023 · 2023
Later among the works it cites.
Speech driven video editing via an audio-conditioned diffusion model
Dan Bigioi, Shubhajit Basak, Michał Stypułkowski, Maciej Zieba, Hugh Jordan, Rachel McDonnell, and Peter Corcoran. 2024 · 2024
Closest in time.
UniCATS: A unified context-aware text-to-speech framework with contextual vq-diffusion and vocoding. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 17924–17932
Chenpeng Du, Yiwei Guo, Feiyu Shen, Zhijun Liu, Zheng Liang, Xie Chen, Shuai Wang, Hui Zhang, and Kai Yu. 2024 · 2024
Closest in time.
GAIA: Zero-shot Talking Avatar Generation
Tianyu He, Junliang Guo, Runyi Yu, Yuchi Wang, Jialiang Zhu, Kaikai An, Leyi Li, Xu Tan, Chunyu Wang, Han Hu, HsiangTao Wu, Sheng Zhao, and Jiang Bian. 2024 · 2024
Closest in time.
DiffDub: Person-Generic Visual Dubbing Using Inpainting Renderer with Diffusion Auto-Encoder. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 3630–3634
Tao Liu, Chenpeng Du, Shuai Fan, Feilong Chen, and Kai Yu. 2024 · 2024
Closest in time.
DiffSpeaker: Speech-Driven 3D Facial Animation with Diffusion Transformer
Zhiyuan Ma, Xiangyu Zhu, Guojun Qi, Chen Qian, Zhaoxiang Zhang, and Zhen Lei. 2024 · 2024
Closest in time.
Diff2lip: Audio conditioned diffusion models for lip-synchronization. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision . 5292–5302
Soumik Mukhopadhyay, Saksham Suri, Ravi Teja Gadde, and Abhinav Shrivastava. 2024 · 2024
Closest in time.
Exploring Phonetic Context-Aware Lip-Sync for Talking Face Generation. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 4325–4329
Se Jin Park, Minsu Kim, Jeongsoo Choi, and Yong Man Ro. 2024 · 2024
Closest in time.
Diffused heads: Diffusion models beat gans on talking-face generation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision . 5091–5100
Michał Stypułkowski, Konstantinos Vougioukas, Sen He, Maciej Zięba, Stavros Petridis, and Maja Pantic. 2024 · 2024
Closest in time.
Linrui Tian, Qi Wang, Bang Zhang, and Liefeng Bo. 2024 · 2024
Closest in time.