Fetching the paper…
Reading the bibliography…
Generating high-fidelity talking head video by fitting with the input audio sequence is a challenging problem that receives considerable attentions recently.
Confusions among visually perceived consonants
Cletus G Fisher · 1968
Earlier work this paper cites.
Video rewrite: driving visual speech with audio
Christoph Bregler, Michele Covell, and Malcolm Slaney · 1997
Earlier work this paper cites.
Multiple view geometry in computer vision
Alex M Andrew · 2001
Earlier work this paper cites.
Trainable videorealistic speech animation
Tony Ezzat, Gadi Geiger, and Tomaso Poggio · 2002
Earlier work this paper cites.
Poisson image editing
Patrick Pérez, Michel Gangnet, and Andrew Blake · 2003
Earlier work this paper cites.
Expressive speech-driven facial animation
Yong Cao, Wen C Tien, Petros Faloutsos, and Frédéric Pighin · 2005
Earlier work this paper cites.
Performance-driven facial animation
Lance Williams · 2006
Earlier work this paper cites.
Real-time vision and speech driven avatars for multimedia applications
Oliver Schreer, Roman Englert, Peter Eisert, and Ralf Tanger · 2008
Earlier work this paper cites.
A variational approach to video registration with subspace constraints
Ravi Garg, Anastasios Roussos, and Lourdes Agapito · 2013
Earlier work this paper cites.
Cross-dataset learning and person-specific normalisation for automatic action unit detection
T. Baltrušaitis, M. Mahmoud, and P. Robinson · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederick P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Real-time expression transfer for facial reenactment
Justus Thies, Michael Zollhöfer, Matthias Nießner, Levi Valgaerts, Marc Stamminger, and Christian Theobalt · 2015
Earlier work this paper cites.
Deep speech 2: End-to-end speech recognition in english and mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, et al · 2016
Earlier work this paper cites.
Out of time: automated lip sync in the wild
J. S. Chung and A. Zisserman · 2016
Earlier work this paper cites.
Jali: an animator-centric viseme model for expressive lip synchronization
Pif Edwards, Chris Landreth, Eugene Fiume, and Karan Singh · 2016
Earlier work this paper cites.
Face2face: Real-time face capture and reenactment of rgb videos
Justus Thies, Michael Zollhofer, Marc Stamminger, Christian Theobalt, and Matthias Nießner · 2016
Earlier work this paper cites.
You said that?
Joon Son Chung, Amir Jamaludin, and Andrew Zisserman · 2017
Earlier work this paper cites.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen · 2017
Earlier work this paper cites.
Speech-driven 3d facial animation with implicit emotional awareness: A deep learning approach
Hai X Pham, Samuel Cheung, and Vladimir Pavlovic · 2017
Earlier work this paper cites.
Synthesizing obama: learning lip sync from audio
Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman · 2017
Earlier work this paper cites.
A deep learning approach for generalized speech animation
Sarah Taylor, Taehwan Kim, Yisong Yue, Moshe Mahler, James Krahe, Anastasio Garcia Rodriguez, Jessica Hodgins, and Iain Matthews · 2017
Cited alongside, same era.
Openface 2.0: Facial behavior analysis toolkit
T. Baltrusaitis, A. Zadeh, Y. C. Lim, and L. Morency · 2018
Cited alongside, same era.
Cnn-based real-time dense face reconstruction with inverse-rendered photo-realistic face images
Yudong Guo, Jianfei Cai, Boyi Jiang, Jianmin Zheng, et al · 2018
Cited alongside, same era.
Deep video portraits
Hyeongwoo Kim, Pablo Garrido, Ayush Tewari, Weipeng Xu, Justus Thies, Matthias Nießner, Patrick Pérez, Christian Richardt, Michael Zollöfer, and Christian Theobalt · 2018
Cited alongside, same era.
X2face: A network for controlling face generation using images, audio, and pose codes
Olivia Wiles, A Sophia Koepke, and Andrew Zisserman · 2018
Cited alongside, same era.
Nerf: Representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng · 2020
Later among the works it cites.
Deformable neural radiance fields
Keunhong Park, Utkarsh Sinha, Jonathan T. Barron, Sofien Bouaziz, Dan B Goldman, Steven M. Seitz, and Ricardo Martin-Brualla · 2020
Later among the works it cites.
Neural voice puppetry: Audio-driven facial reenactment
Justus Thies, Mohamed Elgharib, Ayush Tewari, Christian Theobalt, and Matthias Nießner · 2020
Later among the works it cites.
Mead: A large-scale audio-visual dataset for emotional talking-face generation
Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, and Chen Change Loy · 2020
Later among the works it cites.
Lightweight photometric stereo for facial details recovery
Xueying Wang, Yudong Guo, Bailin Deng, and Juyong Zhang · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yang Zhou, Zhan Xu, Chris Landreth, Evangelos Kalogerakis, Subhransu Maji, and Karan Singh · 2018
Cited alongside, same era.
Hierarchical cross-modal talking face generation with dynamic pixel-wise loss
Lele Chen, Ross K Maddox, Zhiyao Duan, and Chenliang Xu · 2019
Cited alongside, same era.
Capture, learning, and synthesis of 3d speaking styles
Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael J Black · 2019
Cited alongside, same era.
Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set
Yu Deng, Jiaolong Yang, Sicheng Xu, Dong Chen, Yunde Jia, and Xin Tong · 2019
Cited alongside, same era.
Text-based editing of talking-head video
Ohad Fried, Ayush Tewari, Michael Zollhöfer, Adam Finkelstein, Eli Shechtman, Dan B Goldman, Kyle Genova, Zeyu Jin, Christian Theobalt, and Maneesh Agrawala · 2019
Cited alongside, same era.
Neural style-preserving visual dubbing
Hyeongwoo Kim, Mohamed Elgharib, Hans-Peter Zollöfer, Michael Seidel, Thabo Beeler, Christian Richardt, and Christian Theobalt · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Cited alongside, same era.
Ran Yi, Zipeng Ye, Juyong Zhang, Hujun Bao, and Yong-Jin Liu · 2020
Later among the works it cites.
Dynamic neural radiance fields for monocular 4d facial avatar reconstruction
Guy Gafni, Justus Thies, Michael Zollhöfer, and Matthias Nießner · 2021
Closest in time.
Stereopifu: Depth aware clothed human digitization via stereo vision
Yang Hong, Juyong Zhang, Boyi Jiang, Yudong Guo, Ligang Liu, and Hujun Bao · 2021
Closest in time.
Audio-driven emotional video portraits
Xinya Ji, Hang Zhou, Kaisiyuan Wang, Wayne Wu, Chen Change Loy, Xun Cao, and Feng Xu · 2021
Closest in time.
Write-a-speaker: Text-based emotional and rhythmic talking-head generation
Lincheng Li, Suzhen Wang, Zhimeng Zhang, Yu Ding, Yixing Zheng, Xin Yu, and Changjie Fan · 2021
Closest in time.
NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections
Ricardo Martin-Brualla, Noha Radwan, Mehdi S. M. Sajjadi, Jonathan T. Barron, Alexey Dosovitskiy, and Daniel Duckworth · 2021
Closest in time.
Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans
Sida Peng, Yuanqing Zhang, Yinghao Xu, Qianqian Wang, Qing Shuai, Hujun Bao, and Xiaowei Zhou · 2021
Closest in time.
D-nerf: Neural radiance fields for dynamic scenes
Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer · 2021
Closest in time.
Pixel-aligned volumetric avatars
Amit Raj, Michael Zollhofer, Tomas Simon, Jason Saragih, Shunsuke Saito, James Hays, and Stephen Lombardi · 2021
Closest in time.
Audio2head: Audio-driven one-shot talking-head generation with natural head motion
Suzhen Wang, Lincheng Li, Yu Ding, Changjie Fan, and Xin Yu · 2021
Closest in time.
One-shot free-view neural talking-head synthesis for video conferencing
Ting-Chun Wang, Arun Mallya, and Ming-Yu Liu · 2021
Closest in time.
Learning compositional radiance fields of dynamic human heads, June 2021
Ziyan Wang, Timur Bagautdinov, Stephen Lombardi, Tomas Simon, Jason Saragih, Jessica Hodgins, and Michael Zollhofer · 2021
Closest in time.
Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset
Zhimeng Zhang, Lincheng Li, Yu Ding, and Changjie Fan · 2021
Closest in time.
Pose-controllable talking face generation by implicitly modularized audio-visual representation
Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu · 2021
Closest in time.