Fetching the paper…
Reading the bibliography…
Recently, talking-face video generation has received considerable attention.
A Morphable Model for the Synthesis of 3D Faces. In Proceedings of the 26th Annual Conference on Computer Graphics and Interactive Techniques (SIGGRAPH) . 187–194
Volker Blanz and Thomas Vetter. 1999 · 1999
Earlier work this paper cites.
Generative Adversarial Nets. In Advances in Neural Information Processing Systems (NeurIPS) . 2672–2680
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Out of time: automated lip sync in the wild. In Asian conference on computer vision . Springer, 251–263
Joon Son Chung and Andrew Zisserman. 2016 · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
You said that?. In British Machine Vision Conference (BMVC)
Joon Son Chung, Amir Jamaludin, and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen. 2017 · 2017
Earlier work this paper cites.
Realistic dynamic facial textures from a single image using gans. In Proceedings of the IEEE International Conference on Computer Vision . 5429–5438
Kyle Olszewski, Zimo Li, Chao Yang, Yi Zhou, Ronald Yu, Zeng Huang, Sitao Xiang, Shunsuke Saito, Pushmeet Kohli, and Hao Li. 2017 · 2017
Earlier work this paper cites.
Automatic differentiation in PyTorch. In NIPS 2017 Autodiff Workshop: The Future of Gradient-based Machine Learning Software and Techniques
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. 2017 · 2017
Earlier work this paper cites.
Lip Movements Generation at a Glance. In Proceedings of the European Conference on Computer Vision (ECCV) . 538–553
Lele Chen, Zhiheng Li, Ross K. Maddox, Zhiyao Duan, and Chenliang Xu. 2018 · 2018
Earlier work this paper cites.
Stargan: Unified generative adversarial networks for multi-domain image-to-image translation. In Proceedings of the IEEE conference on computer vision and pattern recognition . 8789–8797
Yunjey Choi, Minje Choi, Munyoung Kim, Jung-Woo Ha, Sunghun Kim, and Jaegul Choo. 2018 · 2018
Earlier work this paper cites.
Exprgan: Facial expression editing with controllable expression intensity. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 32
Hui Ding, Kumar Sricharan, and Rama Chellappa. 2018 · 2018
Earlier work this paper cites.
paGAN: real-time avatars using dynamic textures
Koki Nagano, Jaewoo Seo, Jun Xing, Lingyu Wei, Zimo Li, Shunsuke Saito, Aviral Agarwal, Jens Fursund, and Hao Li. 2018 · 2018
Earlier work this paper cites.
An ensemble framework of voice-based emotion recognition system for films and TV programs. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6209–6213
Fei Tao, Gang Liu, and Qingen Zhao. 2018 · 2018
Cited alongside, same era.
Hierarchical Cross-Modal Talking Face Generation with Dynamic Pixel-Wise Loss. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 7832–7841
Lele Chen, Ross K. Maddox, Zhiyao Duan, and Chenliang Xu. 2019 · 2019
Cited alongside, same era.
Capture, learning, and synthesis of 3D speaking styles. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10101–10111
Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael J Black. 2019 · 2019
Cited alongside, same era.
3d guided fine-grained face manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 9821–9830
Zhenglin Geng, Chen Cao, and Sergey Tulyakov. 2019 · 2019
Cited alongside, same era.
State of the art on neural rendering. In Computer Graphics Forum , Vol. 39. Wiley Online Library, 701–727
Ayush Tewari, Ohad Fried, Justus Thies, Vincent Sitzmann, Stephen Lombardi, Kalyan Sunkavalli, Ricardo Martin-Brualla, Tomas Simon, Jason Saragih, Matthias Nießner, et al · 2020
Later among the works it cites.
Neural voice puppetry: Audio-driven facial reenactment. In European Conference on Computer Vision . Springer, 716–731
Justus Thies, Mohamed Elgharib, Ayush Tewari, Christian Theobalt, and Matthias Nießner. 2020 · 2020
Later among the works it cites.
Mead: A large-scale audio-visual dataset for emotional talking-face generation. In European Conference on Computer Vision . Springer, 700–717
Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, and Chen Change Loy. 2020 · 2020
Later among the works it cites.
Photorealistic Audio-driven Video Portraits
Xin Wen, Miao Wang, Christian Richardt, Ze-Yin Chen, and Shi-Min Hu. 2020 · 2020
Later among the works it cites.
Leed: Label-free expression editing via disentanglement. In European Conference on Computer Vision . Springer, 781–798
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 4401–4410
Tero Karras, Samuli Laine, and Timo Aila. 2019 · 2019
Cited alongside, same era.
Real-Time Facial Expression Transformation for Monocular RGB Video. In Computer Graphics Forum , Vol. 38. Wiley Online Library, 470–481
Luming Ma and Zhigang Deng. 2019 · 2019
Cited alongside, same era.
Talking Face Generation by Conditional Recurrent Adversarial Network. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence (IJCAI) . 919–925
Yang Song, Jingwen Zhu, Dawei Li, Andy Wang, and Hairong Qi. 2019 · 2019
Cited alongside, same era.
Deferred neural rendering: Image synthesis using neural textures
Justus Thies, Michael Zollhöfer, and Matthias Nießner. 2019 · 2019
Cited alongside, same era.
Talking Face Generation by Adversarially Disentangled Audio-Visual Representation. In The Thirty-Third AAAI Conference on Artificial Intelligence (AAAI) . 9299–9306
Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang. 2019 · 2019
Cited alongside, same era.
Nerf: Representing scenes as neural radiance fields for view synthesis. In European conference on computer vision . Springer, 405–421
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. 2020 · 2020
Cited alongside, same era.
A Lip Sync Expert Is All You Need for Speech to Lip Generation In the Wild. In The 28th ACM International Conference on Multimedia (MM) . 484–492
K. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, and C. V. Jawahar. 2020 · 2020
Cited alongside, same era.
Accelerating 3D Deep Learning with PyTorch3D
Nikhila Ravi, Jeremy Reizenstein, David Novotny, Taylor Gordon, Wan-Yen Lo, Justin Johnson, and Georgia Gkioxari. 2020 · 2020
Cited alongside, same era.
Rongliang Wu and Shijian Lu. 2020 · 2020
Later among the works it cites.
Audio-driven Talking Face Video Generation with Natural Head Pose
Ran Yi, Zipeng Ye, Juyong Zhang, Hujun Bao, and Yong-Jin Liu. 2020 · 2020
Later among the works it cites.
MakeltTalk: speaker-aware talking-head animation
Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li. 2020 · 2020
Later among the works it cites.
Audio-driven emotional video portraits. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 14080–14089
Xinya Ji, Hang Zhou, Kaisiyuan Wang, Wayne Wu, Chen Change Loy, Xun Cao, and Feng Xu. 2021 · 2021
Later among the works it cites.
Encoding in style: a stylegan encoder for image-to-image translation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 2287–2296
Elad Richardson, Yuval Alaluf, Or Patashnik, Yotam Nitzan, Yaniv Azar, Stav Shapiro, and Daniel Cohen-Or. 2021 · 2021
Later among the works it cites.
Imitating arbitrary talking style for realistic audio-driven talking face synthesis. In Proceedings of the 29th ACM International Conference on Multimedia . 1478–1486
Haozhe Wu, Jia Jia, Haoyu Wang, Yishun Dou, Chao Duan, and Qingshan Deng. 2021 · 2021
Later among the works it cites.
Weakly-Supervised Multi-Face 3D Reconstruction
Jialiang Zhang, Lixiang Lin, Jianke Zhu, and Steven CH Hoi. 2021 · 2021
Later among the works it cites.
Audio-Driven Talking Face Video Generation with Dynamic Convolution Kernels
Zipeng Ye, Mengfei Xia, Ran Yi, Juyong Zhang, Yu-Kun Lai, Xuwei Huang, Guoxin Zhang, and Yong-jin Liu. 2022 · 2022
Closest in time.