Fetching the paper…
Reading the bibliography…
We present a method that generates expressive talking heads from a single facial image with audio as the only input.
Hearing lips and seeing voices
Harry McGurk and John MacDonald. 1976 · 1976
Earlier work this paper cites.
Facial action coding system: a technique for the measurement of facial movement
Paul Ekman and Wallace V Friesen. 1978 · 1978
Earlier work this paper cites.
Improvising linguistic style: Social and affective bases for agent personality. In Proc. IAA
Marilyn A Walker, Janet E Cahn, and Stephen J Whittaker. 1997 · 1997
Earlier work this paper cites.
Voice puppetry. In Proc. SIGGRAPH
Matthew Brand. 1999 · 1999
Earlier work this paper cites.
Dancing-to-music character animation. In Computer Graphics Forum
Takaaki Shiratori, Atsushi Nakazawa, and Katsushi Ikeuchi. 2006 · 2006
Earlier work this paper cites.
Differential representations for mesh processing. In Computer Graphics Forum
Olga Sorkine. 2006 · 2006
Earlier work this paper cites.
Visualizing data using t-SNE
Laurens van der Maaten and Geoffrey Hinton. 2008 · 2008
Earlier work this paper cites.
Generalized-icp.. In Proc. Robotics: science and systems
Aleksandr Segal, Dirk Haehnel, and Sebastian Thrun. 2009 · 2009
Earlier work this paper cites.
Voice transformation: a survey. In Proc. ICASSP
Yannis Stylianou. 2009 · 2009
Earlier work this paper cites.
The artist’s complete guide to facial expression
Gary Faigin. 2012 · 2012
Earlier work this paper cites.
Rigid stabilization of facial expressions
Thabo Beeler and Derek Bradley. 2014 · 2014
Earlier work this paper cites.
Towards end-to-end speech recognition with recurrent neural networks. In Proc. ICML
Alex Graves and Navdeep Jaitly. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition. In Proc. ICLR
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Vdub: Modifying face video of actors for plausible visual alignment to a dubbed audio track. In Computer graphics forum
Pablo Garrido, Levi Valgaerts, Hamid Sarmadi, Ingmar Steiner, Kiran Varanasi, Patrick Perez, and Christian Theobalt. 2015 · 2015
Earlier work this paper cites.
Video-audio driven real-time facial animation
Yilong Liu, Feng Xu, Jinxiang Chai, Xin Tong, Lijuan Wang, and Qiang Huo. 2015 · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation. In Proc. MICCAI
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015 · 2015
Earlier work this paper cites.
JALI: an animator-centric viseme model for expressive lip synchronization
Pif Edwards, Chris Landreth, Eugene Fiume, and Karan Singh. 2016 · 2016
Earlier work this paper cites.
Perceptual losses for real-time style transfer and super-resolution. In Proc. ECCV
Justin Johnson, Alexandre Alahi, and Li Fei-Fei. 2016 · 2016
Earlier work this paper cites.
Face2face: Real-time face capture and reenactment of rgb videos. In Proc. CVPR
Justus Thies, Michael Zollhofer, Marc Stamminger, Christian Theobalt, and Matthias Nießner. 2016 · 2016
Earlier work this paper cites.
Superseded-cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit
Christophe Veaux, Junichi Yamagishi, Kirsten MacDonald, et al · 2016
Earlier work this paper cites.
Bringing portraits to life
Hadar Averbuch-Elor, Daniel Cohen-Or, Johannes Kopf, and Michael F Cohen. 2017 · 2017
Cited alongside, same era.
How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks). In Proc. ICCV
Adrian Bulat and Georgios Tzimiropoulos. 2017 · 2017
Cited alongside, same era.
You said that?. In Proc. BMVC
Joon Son Chung, Amir Jamaludin, and Andrew Zisserman. 2017 · 2017
Cited alongside, same era.
Example-Based Synthesis of Stylized Facial Animations
Jakub Fišer, Ondřej Jamriška, David Simons, Eli Shechtman, Jingwan Lu, Paul Asente, Michal Lukáč, and Daniel Sýkora. 2017 · 2017
Cited alongside, same era.
Image-to-image translation with conditional adversarial networks. In Proc. CVPR
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. 2017 · 2017
Cited alongside, same era.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
Protecting World Leaders Against Deep Fakes. In Proc. CVPRW
Shruti Agarwal, Hany Farid, Yuming Gu, Mingming He, Koki Nagano, and Hao Li. 2019 · 2019
Later among the works it cites.
Hierarchical cross-modal talking face generation with dynamic pixel-wise loss. In Proc. CVPR
Lele Chen, Ross K Maddox, Zhiyao Duan, and Chenliang Xu. 2019 · 2019
Later among the works it cites.
Capture, Learning, and Synthesis of 3D Speaking Styles. In Proc. CVPR
Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael J Black. 2019 · 2019
Later among the works it cites.
Noise-Resilient Training Method for Face Landmark Generation From Speech
Sefik Emre Eskimez, Ross K Maddox, Chenliang Xu, and Zhiyao Duan. 2019 · 2019
Later among the works it cites.
Learning Individual Styles of Conversational Gesture. In Proc. CVPR
Shiry Ginosar, Amir Bar, Gefen Kohavi, Caroline Chan, Andrew Owens, and Jitendra Malik. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen. 2017 · 2017
Cited alongside, same era.
Least squares generative adversarial networks. In Proc. ICCV
Xudong Mao, Qing Li, Haoran Xie, Raymond YK Lau, Zhen Wang, and Stephen Paul Smolley. 2017 · 2017
Cited alongside, same era.
Synthesizing obama: learning lip sync from audio
Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman. 2017 · 2017
Cited alongside, same era.
A deep learning approach for generalized speech animation
Sarah Taylor, Taehwan Kim, Yisong Yue, Moshe Mahler, James Krahe, Anastasio Garcia Rodriguez, Jessica Hodgins, and Iain Matthews. 2017 · 2017
Cited alongside, same era.
Attention is all you need. In Proc. NeurIPS
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Visemenet: Audio-driven animator-centric speech animation
Yang Zhou, Zhan Xu, Chris Landreth, Evangelos Kalogerakis, Subhransu Maji, and Karan Singh. 2018 · 2017
Cited alongside, same era.
Lip movements generation at a glance. In Proc. ECCV
Lele Chen, Zhiheng Li, Ross K Maddox, Zhiyao Duan, and Chenliang Xu. 2018 · 2018
Cited alongside, same era.
FiNet: Compatible and Diverse Fashion Image Inpainting. In Proc. ICCV
Xintong Han, Zuxuan Wu, Weilin Huang, Matthew R Scott, and Larry S Davis. 2019 · 2019
Later among the works it cites.
Neural style-preserving visual dubbing
Hyeongwoo Kim, Mohamed Elgharib, Michael Zollhöfer, Hans-Peter Seidel, Thabo Beeler, Christian Richardt, and Christian Theobalt. 2019 · 2019
Later among the works it cites.
AUTOVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss. In Proc. ICML . 5210–5219
Kaizhi Qian, Yang Zhang, Shiyu Chang, Xuesong Yang, and Mark Hasegawa-Johnson. 2019 · 2019
Later among the works it cites.
First Order Motion Model for Image Animation. In Conference on Neural Information Processing Systems (NeurIPS)
Aliaksandr Siarohin, Stéphane Lathuilière, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. 2019 · 2019
Later among the works it cites.
Talking Face Generation by Conditional Recurrent Adversarial Network. In Proc. IJCAI
Yang Song, Jingwen Zhu, Dawei Li, Andy Wang, and Hairong Qi. 2019 · 2019
Later among the works it cites.
Realistic Speech-Driven Facial Animation with GANs
Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic. 2019 · 2019
Later among the works it cites.
Photo wake-up: 3d character animation from a single photo. In Proc. CVPR
Chung-Yi Weng, Brian Curless, and Ira Kemelmacher-Shlizerman. 2019 · 2019
Later among the works it cites.
The face of art: landmark detection and geometric style in portraits
Jordan Yaniv, Yael Newman, and Ariel Shamir. 2019 · 2019
Later among the works it cites.
Few-Shot Adversarial Learning of Realistic Neural Talking Head Models. In Proc. ICCV
Egor Zakharov, Aliaksandra Shysheya, Egor Burkov, and Victor Lempitsky. 2019 · 2019
Later among the works it cites.
Talking face generation by adversarially disentangled audio-visual representation. In Proc. AAAI
Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang. 2019 · 2019
Later among the works it cites.
Foley Music: Learning to Generate Music from Videos
Chuang Gan, Deng Huang, Peihao Chen, Joshua B Tenenbaum, and Antonio Torralba. 2020a · 2020
Closest in time.
Text to Speech converter
Notevibes. 2020 · 2020
Closest in time.
Background Matting: The World is Your Green Screen. In Proc. CVPR
Soumyadip Sengupta, Vivek Jayaram, Brian Curless, Steve Seitz, and Ira Kemelmacher-Shlizerman. 2020 · 2020
Closest in time.
Neural Voice Puppetry: Audio-driven Facial Reenactment. In Proc. CVPR, to appear
Justus Thies, Mohamed Elgharib, Ayush Tewari, Christian Theobalt, and Matthias Nießner. 2020 · 2020
Closest in time.