Fetching the paper…
Reading the bibliography…
While recent advances in deep neural networks have made it possible to render high-quality images, generating photo-realistic and personalized talking head remains challenging.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Real-time vision and speech driven avatars for multimedia applications
Oliver Schreer, Roman Englert, Peter Eisert, and Ralf Tanger · 2008
Earlier work this paper cites.
Image quality metrics: PSNR vs. SSIM
Alain Horé and Djemel Ziou · 2010
Earlier work this paper cites.
Synthesizing photo-real talking head via trajectory-guided sample selection
Lijuan Wang, Xiaojun Qian, Wei Han, and Frank K. Soong · 2010
Earlier work this paper cites.
On the properties of mean opinion scores for quality of experience management
Jie Xu, Liyuan Xing, Andrew Perkis, and Yuming Jiang · 2011
Earlier work this paper cites.
High quality lip-sync animation for 3d photo-realistic talking head
Lijuan Wang, Wei Han, and Frank K Soong · 2012
Earlier work this paper cites.
An expressive text-driven 3d talking head
Robert Anderson, Björn Stenger, Vincent Wan, and Roberto Cipolla · 2013
Earlier work this paper cites.
Facewarehouse: A 3d facial expression database for visual computing
Chen Cao, Yanlin Weng, Shun Zhou, Yiying Tong, and Kun Zhou · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Photo-real talking head with deep bidirectional lstm
Bo Fan, Lijuan Wang, Frank K Soong, and Lei Xie · 2015
Earlier work this paper cites.
Talking heads synthesis from audio with deep neural networks
Taiki Shimba, Ryuhei Sakurai, Hirotake Yamazoe, and Joo-Ho Lee · 2015
Earlier work this paper cites.
Deep speech 2: End-to-end speech recognition in english and mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al · 2016
Earlier work this paper cites.
Out of time: Automated lip sync in the wild
Joon Son Chung and Andrew Zisserman · 2016
Earlier work this paper cites.
Jali: an animator-centric viseme model for expressive lip synchronization
Pif Edwards, Chris Landreth, Eugene Fiume, and Karan Singh · 2016
Earlier work this paper cites.
A deep bidirectional lstm approach for video-realistic talking head
Bo Fan, Lei Xie, Shan Yang, Lijuan Wang, and Frank K Soong · 2016
Earlier work this paper cites.
Lip reading sentences in the wild
Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman · 2017
Earlier work this paper cites.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Gaussian process prior variational autoencoders
Francesco Paolo Casale, Adrian V. Dalca, Luca Saglietti, Jennifer Listgarten, and Nicoló Fusi · 2018
Earlier work this paper cites.
Lip movements generation at a glance
Lele Chen, Zhiheng Li, Ross K. Maddox, Zhiyao Duan, and Chenliang Xu · 2018
Earlier work this paper cites.
Cnn-based real-time dense face reconstruction with inverse-rendered photo-realistic face images
Yudong Guo, Jianfei Cai, Boyi Jiang, Jianmin Zheng, et al · 2018
Cited alongside, same era.
Seeing voices and hearing faces: Cross-modal biometric matching
Arsha Nagrani, Samuel Albanie, and Andrew Zisserman · 2018
Cited alongside, same era.
Ganimation: Anatomically-aware facial animation from a single image
Albert Pumarola, Antonio Agudo, Aleix M. Martínez, Alberto Sanfeliu, and Francesc Moreno-Noguer · 2018
Cited alongside, same era.
Visemenet: Audio-driven animator-centric speech animation
Yang Zhou, Zhan Xu, Chris Landreth, Evangelos Kalogerakis, Subhransu Maji, and Karan Singh · 2018
Cited alongside, same era.
Hierarchical cross-modal talking face generation with dynamic pixel-wise loss
Lele Chen, Ross K. Maddox, Zhiyao Duan, and Chenliang Xu · 2019
Cited alongside, same era.
Capture, learning, and synthesis of 3d speaking styles
Nerf: Representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng · 2020
Later among the works it cites.
A lip sync expert is all you need for speech to lip generation in the wild
KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar · 2020
Later among the works it cites.
Neural voice puppetry: Audio-driven facial reenactment
Justus Thies, Mohamed Elgharib, Ayush Tewari, Christian Theobalt, and Matthias Nießner · 2020
Later among the works it cites.
Mead: A large-scale audio-visual dataset for emotional talking-face generation
Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, and Chen Change Loy · 2020
Later among the works it cites.
Photorealistic audio-driven video portraits
Xin Wen, Miao Wang, Christian Richardt, Ze-Yin Chen, and Shi-Min Hu · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael J. Black · 2019
Cited alongside, same era.
Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set
Yu Deng, Jiaolong Yang, Sicheng Xu, Dong Chen, Yunde Jia, and Xin Tong · 2019
Cited alongside, same era.
Towards automatic face-to-face translation
Prajwal KR, Rudrabha Mukhopadhyay, Jerin Philip, Abhishek Jha, Vinay Namboodiri, and CV Jawahar · 2019
Cited alongside, same era.
Neural volumes: learning dynamic renderable volumes from images
Stephen Lombardi, Tomas Simon, Jason M. Saragih, Gabriel Schwartz, Andreas M. Lehrmann, and Yaser Sheikh · 2019
Cited alongside, same era.
Combining 3d morphable models: A large scale face-and-head model
Stylianos Ploumpis, Haoyang Wang, Nick E. Pears, William A. P. Smith, and Stefanos Zafeiriou · 2019
Cited alongside, same era.
Autovc: Zero-shot voice style transfer with only autoencoder loss
Kaizhi Qian, Yang Zhang, Shiyu Chang, Xuesong Yang, and Mark Hasegawa-Johnson · 2019
Cited alongside, same era.
Speech-driven expressive talking lips with conditional sequential generative adversarial networks
Najmeh Sadoughi and Carlos Busso · 2019
Cited alongside, same era.
Ran Yi, Zipeng Ye, Juyong Zhang, Hujun Bao, and Yong-Jin Liu · 2020
Later among the works it cites.
Makelttalk: speaker-aware talking-head animation
Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li · 2020
Later among the works it cites.
Arbitrary talking face generation via attentional audio-visual coherence learning
Hao Zhu, Huaibo Huang, Yi Li, Aihua Zheng, and Ran He · 2020
Later among the works it cites.
Dynamic neural radiance fields for monocular 4d facial avatar reconstruction
Guy Gafni, Justus Thies, Michael Zollhofer, and Matthias Nießner · 2021
Later among the works it cites.
Ad-nerf: Audio driven neural radiance fields for talking head synthesis
Yudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu, Hujun Bao, and Juyong Zhang · 2021
Later among the works it cites.
Neural lumigraph rendering
Petr Kellnhofer, Lars C Jebe, Andrew Jones, Ryan Spicer, Kari Pulli, and Gordon Wetzstein · 2021
Later among the works it cites.
Lipsync3d: Data-efficient learning of personalized 3d talking faces from video using pose and lighting normalization
Avisek Lahiri, Vivek Kwatra, Christian Frueh, John Lewis, and Chris Bregler · 2021
Later among the works it cites.
Live Speech Portraits: Real-time photorealistic talking-head animation
Yuanxun Lu, Jinxiang Chai, and Xun Cao · 2021
Later among the works it cites.
D-nerf: Neural radiance fields for dynamic scenes
Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer · 2021
Later among the works it cites.
Anyonenet: Synchronized speech and talking head generation for arbitrary person
Xinsheng Wang, Qicong Xie, Jihua Zhu, Lei Xie, et al · 2021
Later among the works it cites.
Imitating arbitrary talking style for realistic audio-driven talking face synthesis
Haozhe Wu, Jia Jia, Haoyu Wang, Yishun Dou, Chao Duan, and Qingshan Deng · 2021
Later among the works it cites.
Facial: Synthesizing dynamic talking face with implicit attribute learning
Chenxu Zhang, Yifan Zhao, Yifei Huang, Ming Zeng, Saifeng Ni, Madhukar Budagavi, and Xiaohu Guo · 2021
Later among the works it cites.
Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset
Zhimeng Zhang, Lincheng Li, Yu Ding, and Changjie Fan · 2021
Later among the works it cites.
Pose-controllable talking face generation by implicitly modularized audio-visual representation
Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu · 2021
Later among the works it cites.