Fetching the paper…
Reading the bibliography…
We propose StyleTalker, a novel audio-driven talking head generation model that can synthesize a video of a talking person from a single reference image with accurately audio-synced lip shapes, realistic head poses, and eye blinks.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Learning phrase representations using rnn encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling · 2014
Earlier work this paper cites.
Variational inference with normalizing flows
Danilo Jimenez Rezende and Shakir Mohamed · 2015
Earlier work this paper cites.
Out of time: Automated lip sync in the wild
Joon Son Chung and Andrew Zisserman · 2016
Earlier work this paper cites.
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie · 2017
Earlier work this paper cites.
Synthesizing obama: learning lip sync from audio
Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman · 2017
Earlier work this paper cites.
Voxceleb2: Deep speaker recognition
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang · 2018
Earlier work this paper cites.
Image2stylegan: How to embed images into the stylegan latent space?
Rameen Abdal, Yipeng Qin, and Peter Wonka · 2019
Earlier work this paper cites.
Hierarchical cross-modal talking face generation with dynamic pixel-wise loss
Lele Chen, Ross K Maddox, Zhiyao Duan, and Chenliang Xu · 2019
Earlier work this paper cites.
Arcface: Additive angular margin loss for deep face recognition
Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou · 2019
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
Tero Karras, Samuli Laine, and Timo Aila · 2019
Cited alongside, same era.
Towards automatic face-to-face translation
Prajwal KR, Rudrabha Mukhopadhyay, Jerin Philip, Abhishek Jha, Vinay Namboodiri, and CV Jawahar · 2019
Cited alongside, same era.
First order motion model for image animation
Aliaksandr Siarohin, Stéphane Lathuilière, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe · 2019
Cited alongside, same era.
Talking face generation by adversarially disentangled audio-visual representation
Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang · 2019
Cited alongside, same era.
Image2stylegan++: How to edit the embedded images?
Rameen Abdal, Yipeng Qin, and Peter Wonka · 2020
Cited alongside, same era.
Talking-head generation with rhythmic head motion
Lele Chen, Guofeng Cui, Celong Liu, Zhong Li, Ziyi Kou, Yi Xu, and Chenliang Xu · 2020
Neural voice puppetry: Audio-driven facial reenactment
Justus Thies, Mohamed Elgharib, Ayush Tewari, Christian Theobalt, and Matthias Nießner · 2020
Later among the works it cites.
Realistic speech-driven facial animation with gans
Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic · 2020
Later among the works it cites.
Makelttalk: speaker-aware talking-head animation
Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li · 2020
Later among the works it cites.
Alias-free generative adversarial networks
Tero Karras, Miika Aittala, Samuli Laine, Erik Härkönen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila · 2021
Later among the works it cites.
Lipsync3d: Data-efficient learning of personalized 3d talking faces from video using pose and lighting normalization
Avisek Lahiri, Vivek Kwatra, Christian Frueh, John Lewis, and Chris Bregler · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Speech-driven facial animation using cascaded gans for learning of motion and texture
Dipanjan Das, Sandika Biswas, Sanjana Sinha, and Brojeshwar Bhowmick · 2020
Cited alongside, same era.
Squeeze-and-excitation networks
Jie Hu, Li Shen, Samuel Albanie, Gang Sun, and Enhua Wu · 2020
Cited alongside, same era.
Contragan: Contrastive learning for conditional image generation
Mingu Kang and Jaesik Park · 2020
Cited alongside, same era.
Analyzing and improving the image quality of stylegan
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila · 2020
Cited alongside, same era.
A lip sync expert is all you need for speech to lip generation in the wild
KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar · 2020
Cited alongside, same era.
Disentanglement by nonlinear ica with general incompressible-flow networks (gin)
Peter Sorrenson, Carsten Rother, and U. Köthe · 2020
Cited alongside, same era.
Rayhane Mama, Marc S Tyndel, Hashiam Kadhim, Cole Clifford, and Ragavan Thurairatnam · 2021
Later among the works it cites.
Encoding in style: a stylegan encoder for image-to-image translation
Elad Richardson, Yuval Alaluf, Or Patashnik, Yotam Nitzan, Yaniv Azar, Stav Shapiro, and Daniel Cohen-Or · 2021
Later among the works it cites.
Designing an encoder for stylegan image manipulation
Omer Tov, Yuval Alaluf, Yotam Nitzan, Or Patashnik, and Daniel Cohen-Or · 2021
Later among the works it cites.
Flowtron: an autoregressive flow-based generative network for text-to-speech synthesis
Rafael Valle, Kevin J. Shih, Ryan J. Prenger, and Bryan Catanzaro · 2021
Later among the works it cites.
Audio2head: Audio-driven one-shot talking-head generation with natural head motion
Suzhen Wang, Lincheng Li, Yu Ding, Changjie Fan, and Xin Yu · 2021
Later among the works it cites.
Cross-modal contrastive learning for text-to-image generation
Han Zhang, Jing Yu Koh, Jason Baldridge, Honglak Lee, and Yinfei Yang · 2021
Later among the works it cites.
Pose-controllable talking face generation by implicitly modularized audio-visual representation
Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu · 2021
Later among the works it cites.