Fetching the paper…
Reading the bibliography…
Although significant progress has been made to audio-driven talking face generation, existing methods either neglect facial emotion or cannot be applied to arbitrary subjects.
Write-a-speaker: Text-based Emotional and Rhythmic Talking-head Generation. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. 1911–1920
Lincheng Li, Suzhen Wang, Zhimeng Zhang, Yu Ding, Yixing Zheng, Xin Yu, and Changjie Fan. 2021 · 1920
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Voice puppetry. In Proceedings of the 26th annual conference on Computer graphics and interactive techniques . 21–28
Matthew Brand. 1999 · 1999
Earlier work this paper cites.
Mel frequency cepstral coefficients for music modeling. In In International Symposium on Music Information Retrieval . Citeseer
Beth Logan. 2000 · 2000
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. 2004 · 2004
Earlier work this paper cites.
High quality lip-sync animation for 3D photo-realistic talking head. In 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 4529–4532
Lijuan Wang, Wei Han, and Frank K Soong. 2012 · 2012
Earlier work this paper cites.
An expressive text-driven 3D talking head
Robert Anderson, Björn Stenger, Vincent Wan, and Roberto Cipolla. 2013 · 2013
Earlier work this paper cites.
Crema-d: Crowd-sourced emotional multimodal actors dataset
Houwei Cao, David G Cooper, Michael K Keutmann, Ruben C Gur, Ani Nenkova, and Ragini Verma. 2014 · 2014
Earlier work this paper cites.
The Chicago face database: A free stimulus set of faces and norming data
Debbie S Ma, Joshua Correll, and Bernd Wittenbrink. 2015 · 2015
Earlier work this paper cites.
JALI: an animator-centric viseme model for expressive lip synchronization
Pif Edwards, Chris Landreth, Eugene Fiume, and Karan Singh. 2016 · 2016
Earlier work this paper cites.
Perceptual losses for real-time style transfer and super-resolution. In European conference on computer vision . Springer, 694–711
Justin Johnson, Alexandre Alahi, and Li Fei-Fei. 2016 · 2016
Earlier work this paper cites.
Ambient sound provides supervision for visual learning. In European conference on computer vision . Springer, 801–816
Andrew Owens, Jiajun Wu, Josh H McDermott, William T Freeman, and Antonio Torralba. 2016 · 2016
Earlier work this paper cites.
Face2face: Real-time face capture and reenactment of rgb videos. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2387–2395
Justus Thies, Michael Zollhofer, Marc Stamminger, Christian Theobalt, and Matthias Nießner. 2016 · 2016
Earlier work this paper cites.
Look, listen and learn. In Proceedings of the IEEE International Conference on Computer Vision . 609–617
Relja Arandjelovic and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
Joon Son Chung, Amir Jamaludin, and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
Synthesizing normalized faces from facial identity features. In Proceedings of the IEEE conference on computer vision and pattern recognition . 3703–3712
Forrester Cole, David Belanger, Dilip Krishnan, Aaron Sarna, Inbar Mosseri, and William T Freeman. 2017 · 2017
Earlier work this paper cites.
Arbitrary style transfer in real-time with adaptive instance normalization. In Proceedings of the IEEE International Conference on Computer Vision . 1501–1510
Xun Huang and Serge Belongie. 2017 · 2017
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition . 1125–1134
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. 2017 · 2017
Earlier work this paper cites.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen. 2017 · 2017
Earlier work this paper cites.
Synthesizing obama: learning lip sync from audio
Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman. 2017 · 2017
Earlier work this paper cites.
Deepfake video detection using recurrent neural networks. In 2018 15th IEEE international conference on advanced video and signal based surveillance (AVSS) . IEEE, 1–6
David Güera and Edward J Delp. 2018 · 2018
Cited alongside, same era.
Deep video portraits
Hyeongwoo Kim, Pablo Garrido, Ayush Tewari, Weipeng Xu, Justus Thies, Matthias Niessner, Patrick Pérez, Christian Richardt, Michael Zollhöfer, and Christian Theobalt. 2018 · 2018
Cited alongside, same era.
Cooperative learning of audio and video models from self-supervised synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani. 2018 · 2018
Cited alongside, same era.
The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in North American English
Steven R Livingstone and Frank A Russo. 2018 · 2018
Cited alongside, same era.
End-to-end speech-driven facial animation with temporal gans
Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic. 2018 · 2018
Towards fast, accurate and stable 3d dense face alignment. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XIX 16 . Springer, 152–168
Jianzhu Guo, Xiangyu Zhu, Yang Yang, Fan Yang, Zhen Lei, and Stan Z Li. 2020 · 2020
Later among the works it cites.
Learning Identity-Invariant Motion Representations for Cross-ID Face Reenactment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 7084–7092
Po-Hsiang Huang, Fu-En Yang, and Yu-Chiang Frank Wang. 2020 · 2020
Later among the works it cites.
Face x-ray for more general face forgery detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 5001–5010
Lingzhi Li, Jianmin Bao, Ting Zhang, Hao Yang, Dong Chen, Fang Wen, and Baining Guo. 2020 · 2020
Later among the works it cites.
Nerf: Representing scenes as neural radiance fields for view synthesis. In European conference on computer vision . Springer, 405–421
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reenactgan: Learning to reenact faces via boundary transfer. In Proceedings of the European conference on computer vision (ECCV) . 603–619
Wayne Wu, Yunxuan Zhang, Cheng Li, Chen Qian, and Chen Change Loy. 2018 · 2018
Cited alongside, same era.
Visemenet: Audio-driven animator-centric speech animation
Yang Zhou, Zhan Xu, Chris Landreth, Evangelos Kalogerakis, Subhransu Maji, and Karan Singh. 2018 · 2018
Cited alongside, same era.
Michael Zollhöfer, Justus Thies, Pablo Garrido, Derek Bradley, Thabo Beeler, Patrick Pérez, Marc Stamminger, Matthias Nießner, and Christian Theobalt. 2018 · 2018
Cited alongside, same era.
Text-based editing of talking-head video
Ohad Fried, Ayush Tewari, Michael Zollhöfer, Adam Finkelstein, Eli Shechtman, Dan B Goldman, Kyle Genova, Zeyu Jin, Christian Theobalt, and Maneesh Agrawala. 2019 · 2019
Cited alongside, same era.
Neural style-preserving visual dubbing
Hyeongwoo Kim, Mohamed Elgharib, Michael Zollhöfer, Hans-Peter Seidel, Thabo Beeler, Christian Richardt, and Christian Theobalt. 2019 · 2019
Cited alongside, same era.
Frame attention networks for facial expression recognition in videos. In 2019 IEEE International Conference on Image Processing (ICIP) . IEEE, 3866–3870
Debin Meng, Xiaojiang Peng, Kai Wang, and Yu Qiao. 2019 · 2019
Cited alongside, same era.
Faceforensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 1–11
Andreas Rossler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. 2019 · 2019
Cited alongside, same era.
Animating face using disentangled audio representations. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision . 3290–3298
Gaurav Mittal and Baoyuan Wang. 2020 · 2020
Later among the works it cites.
A lip sync expert is all you need for speech to lip generation in the wild. In Proceedings of the 28th ACM International Conference on Multimedia . 484–492
KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar. 2020 · 2020
Later among the works it cites.
Neural voice puppetry: Audio-driven facial reenactment. In European Conference on Computer Vision . Springer, 716–731
Justus Thies, Mohamed Elgharib, Ayush Tewari, Christian Theobalt, and Matthias Nießner. 2020 · 2020
Later among the works it cites.
Realistic speech-driven facial animation with gans
Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic. 2020 · 2020
Later among the works it cites.
Mesh guided one-shot face reenactment using graph convolutional networks. In Proceedings of the 28th ACM International Conference on Multimedia . 1773–1781
Guangming Yao, Yi Yuan, Tianjia Shao, and Kun Zhou. 2020 · 2020
Later among the works it cites.
Fast bi-layer neural synthesis of one-shot realistic head avatars. In European Conference on Computer Vision . Springer, 524–540
Egor Zakharov, Aleksei Ivakhnenko, Aliaksandra Shysheya, and Victor Lempitsky. 2020 · 2020
Later among the works it cites.
Freenet: Multi-identity face reenactment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5326–5335
Jiangning Zhang, Xianfang Zeng, Mengmeng Wang, Yusu Pan, Liang Liu, Yong Liu, Yu Ding, and Changjie Fan. 2020 · 2020
Later among the works it cites.
MakeltTalk: speaker-aware talking-head animation
Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li. 2020 · 2020
Later among the works it cites.
Audio-driven emotional video portraits. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 14080–14089
Xinya Ji, Hang Zhou, Kaisiyuan Wang, Wayne Wu, Chen Change Loy, Xun Cao, and Feng Xu. 2021 · 2021
Later among the works it cites.
Live Speech Portraits: Real-Time Photorealistic Talking-Head Animation
Yuanxun Lu, Jinxiang Chai, and Xun Cao. 2021 · 2021
Later among the works it cites.
Audio-and gaze-driven facial animation of codec avatars. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision . 41–50
Alexander Richard, Colin Lea, Shugao Ma, Jurgen Gall, Fernando De la Torre, and Yaser Sheikh. 2021 · 2021
Later among the works it cites.
FACEGAN: Facial Attribute Controllable rEenactment GAN. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision . 1329–1338
Soumya Tripathy, Juho Kannala, and Esa Rahtu. 2021 · 2021
Later among the works it cites.
Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion
Suzhen Wang, Lincheng Li, Yu Ding, Changjie Fan, and Xin Yu. 2021a · 2021
Later among the works it cites.
Iterative text-based editing of talking-heads using neural retargeting
Xinwei Yao, Ohad Fried, Kayvon Fatahalian, and Maneesh Agrawala. 2021 · 2021
Later among the works it cites.
Pose-controllable talking face generation by implicitly modularized audio-visual representation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 4176–4186
Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu. 2021 · 2021
Later among the works it cites.
Joint Audio-Visual Deepfake Detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 14800–14809
Yipin Zhou and Ser-Nam Lim. 2021 · 2021
Later among the works it cites.