Fetching the paper…
Reading the bibliography…
In this paper, we introduce a simple and novel framework for one-shot audio-driven talking head generation.
Dlib-ml: A machine learning toolkit
Davis E. King · 2009
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Face2face: Real-time face capture and reenactment of rgb videos
Justus Thies, Michael Zollhofer, Marc Stamminger, Christian Theobalt, and Matthias Nießner · 2016
Earlier work this paper cites.
How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks)
Adrian Bulat and Georgios Tzimiropoulos · 2017
Earlier work this paper cites.
Aviv Gabbay, Asaph Shamir, and Shmuel Peleg · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Voxceleb: a large-scale speaker identification dataset
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2017
Earlier work this paper cites.
Voxceleb2: Deep speaker recognition
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Deep video portraits
Hyeongwoo Kim, Pablo Garrido, Ayush Tewari, Weipeng Xu, Justus Thies, Matthias Nießner, Patrick Pérez, Christian Richardt, Michael Zollhöfer, and Christian Theobalt · 2018
Earlier work this paper cites.
X2face: A network for controlling face generation by using images, audio, and pose codes
Andrew Zisserman Olivia Wiles, A. Sophia Koepke · 2018
Earlier work this paper cites.
Mobilenetv2: Inverted residuals and linear bottlenecks
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen · 2018
Earlier work this paper cites.
Perfect match: Improved cross-modal embeddings for audio-visual synchronisation
Soo-Whan Chung, Joon Son Chung, and Hong-Goo Kang · 2019
Earlier work this paper cites.
Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set
Yu Deng, Jiaolong Yang, Sicheng Xu, Dong Chen, Yunde Jia, and Xin Tong · 2019
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
Tero Karras, Samuli Laine, and Timo Aila · 2019
Earlier work this paper cites.
Yonglong Tian, Dilip Krishnan, and Phillip Isola · 2019
Earlier work this paper cites.
Realistic speech-driven facial animation with gans
Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic · 2019
Earlier work this paper cites.
Few-shot adversarial learning of realistic neural talking head models
Egor Zakharov, Aliaksandra Shysheya, Egor Burkov, and Victor Lempitsky · 2019
Earlier work this paper cites.
Talking face generation by adversarially disentangled audio-visual representation
Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang · 2019
Earlier work this paper cites.
Self-Supervised MultiModal Versatile Networks
Jean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelović, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira, Sander Dieleman, and Andrew Zisserman · 2020
Cited alongside, same era.
Neural head reenactment with latent pose descriptors
Egor Burkov, Igor Pasechnik, Artur Grigorev, and Victor Lempitsky · 2020
Cited alongside, same era.
In defence of metric learning for speaker recognition
Joon Son Chung, Jaesung Huh, Seongkyu Mun, Minjae Lee, Hee Soo Heo, Soyeon Choe, Chiheon Ham, Sunghwan Jung, Bong-Jin Lee, and Icksang Han · 2020
Cited alongside, same era.
Generative adversarial networks
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Audio2head: Audio-driven one-shot talking-head generation with natural head motion
Suzhen Wang, Lincheng Li, Yu Ding, Changjie Fan, and Xin Yu · 2021
Later among the works it cites.
Heatmap regression via randomized rounding
Baosheng Yu and Dacheng Tao · 2021
Later among the works it cites.
Facial: Synthesizing dynamic talking face with implicit attribute learning
Chenxu Zhang, Yifan Zhao, Yifei Huang, Ming Zeng, Saifeng Ni, Madhukar Budagavi, and Xiaohu Guo · 2021
Later among the works it cites.
Pose-controllable talking face generation by implicitly modularized audio-visual representation
Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning identity-invariant motion representations for cross-id face reenactment
Po-Hsiang Huang, Fu-En Yang, and Yu-Chiang Frank Wang · 2020
Cited alongside, same era.
Analyzing and improving the image quality of stylegan
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila · 2020
Cited alongside, same era.
Diffwave: A versatile diffusion model for audio synthesis
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro · 2020
Cited alongside, same era.
Animating face using disentangled audio representations
Gaurav Mittal and Baoyuan Wang · 2020
Cited alongside, same era.
Disentangled speech embeddings using cross-modal self-supervision, 2020
Arsha Nagrani, Joon son Chung, Samuel Albanie, and Andrew Zisserman · 2020
Cited alongside, same era.
A lip sync expert is all you need for speech to lip generation in the wild
KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar · 2020
Cited alongside, same era.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2020
Cited alongside, same era.
L2cs-net: Fine-grained gaze estimation in unconstrained environments
Ahmed A Abdelrahman, Thorsten Hempel, Aly Khalifa, and Ayoub Al-Hamadi · 2022
Closest in time.
Talking head from speech audio using a pre-trained image generator
Mohammed M Alghamdi, He Wang, Andrew J Bulpitt, and David C Hogg · 2022
Closest in time.
Finding directions in gan’s latent space for neural face reenactment
Stella Bounareli, Vasileios Argyriou, and Georgios Tzimiropoulos · 2022
Closest in time.
Versatile multi-modal pre-training for human-centric perception
Fangzhou Hong, Liang Pan, Zhongang Cai, and Ziwei Liu · 2022
Closest in time.
Depth-aware generative adversarial network for talking head video generation
Fa-Ting Hong, Longhao Zhang, Li Shen, and Dan Xu · 2022
Closest in time.
Dual-generator face reenactment
Gee-Sern Hsu, Chun-Hung Tsai, and Hung-Yi Wu · 2022
Closest in time.
Eamm: One-shot emotional talking face via audio-based emotion-aware motion model
Xinya Ji, Hang Zhou, Kaisiyuan Wang, Qianyi Wu, Wayne Wu, Feng Xu, and Xun Cao · 2022
Closest in time.
Vocalist: An audio-visual synchronisation model for lips and voices
Venkatesh S Kadandale, Juan F Montesinos, and Gloria Haro · 2022
Closest in time.
Expressive talking head generation with granular audio-visual control
Borong Liang, Yan Pan, Zhizhi Guo, Hang Zhou, Zhibin Hong, Xiaoguang Han, Junyu Han, Jingtuo Liu, Errui Ding, and Jingdong Wang · 2022
Closest in time.
Styletalker: One-shot style-based audio-driven talking head video generation
Dongchan Min, Minyoung Song, and Sung Ju Hwang · 2022
Closest in time.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Closest in time.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Closest in time.
Learning dynamic facial radiance fields for few-shot talking head synthesis
Shuai Shen, Wanhua Li, Zheng Zhu, Yueqi Duan, Jie Zhou, and Jiwen Lu · 2022
Closest in time.
Make-a-video: Text-to-video generation without text-video data
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al · 2022
Closest in time.
Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Daniel Cohen-Or, and Amit H Bermano · 2022
Closest in time.
Bevt: Bert pretraining of video transformers
Rui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Yu-Gang Jiang, Luowei Zhou, and Lu Yuan · 2022
Closest in time.
Image as a foreign language: Beit pretraining for all vision and vision-language tasks
Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, and Furu Wei · 2022
Closest in time.
Dfa-nerf: Personalized talking head generation via disentangled face attributes neural rendering
Shunyu Yao, RuiZhe Zhong, Yichao Yan, Guangtao Zhai, and Xiaokang Yang · 2022
Closest in time.