Fetching the paper…
Reading the bibliography…
We present a novel one-shot talking head synthesis method that achieves disentangled and fine-grained control over lip motion, eye gaze&blink, head pose, and emotional expression.
Video rewrite: Driving visual speech with audio
Christoph Bregler, Michele Covell, and Malcolm Slaney · 1997
Earlier work this paper cites.
Voice puppetry
Matthew Brand · 1999
Earlier work this paper cites.
A 3d face model for pose and illumination invariant face recognition
Pascal Paysan, Reinhard Knothe, Brian Amberg, Sami Romdhani, and Thomas Vetter · 2009
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
Out of time: automated lip sync in the wild
Joon Son Chung and Andrew Zisserman · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
beta-vae: Learning basic visual concepts with a constrained variational framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner · 2016
Earlier work this paper cites.
Face2face: Real-time face capture and reenactment of rgb videos
Justus Thies, Michael Zollhofer, Marc Stamminger, Christian Theobalt, and Matthias Nießner · 2016
Earlier work this paper cites.
How far are we from solving the 2d & 3d face alignment problem?(and a dataset of 230,000 3d facial landmarks)
Adrian Bulat and Georgios Tzimiropoulos · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Affectnet: A database for facial expression, valence, and arousal computing in the wild
Ali Mollahosseini, Behzad Hasani, and Mohammad H Mahoor · 2017
Earlier work this paper cites.
Synthesizing obama: learning lip sync from audio
Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman · 2017
Earlier work this paper cites.
Lip movements generation at a glance
Lele Chen, Zhiheng Li, Ross K Maddox, Zhiyao Duan, and Chenliang Xu · 2018
Earlier work this paper cites.
Isolating sources of disentanglement in variational autoencoders
Ricky TQ Chen, Xuechen Li, Roger B Grosse, and David K Duvenaud · 2018
Earlier work this paper cites.
Voxceleb2: Deep speaker recognition
J. S. Chung, A. Nagrani, and A. Zisserman · 2018
Earlier work this paper cites.
Semantically decomposing the latent spaces of generative adversarial networks
Chris Donahue, Zachary C Lipton, Akshay Balsubramani, and Julian McAuley · 2018
Earlier work this paper cites.
Rt-gene: Real-time eye gaze estimation in natural environments
Tobias Fischer, Hyung Jin Chang, and Yiannis Demiris · 2018
Earlier work this paper cites.
Deep video portraits
Hyeongwoo Kim, Pablo Garrido, Ayush Tewari, Weipeng Xu, Justus Thies, Matthias Niessner, Patrick Pérez, Christian Richardt, Michael Zollhöfer, and Christian Theobalt · 2018
Earlier work this paper cites.
Disentangling by factorising
Hyunjik Kim and Andriy Mnih · 2018
Earlier work this paper cites.
X2face: A network for controlling face generation by using images, audio, and pose codes
Andrew Zisserman Olivia Wiles, A. Sophia Koepke · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Ganimation: Anatomically-aware facial animation from a single image
Albert Pumarola, Antonio Agudo, Aleix M Martinez, Alberto Sanfeliu, and Francesc Moreno-Noguer · 2018
Earlier work this paper cites.
Reenactgan: Learning to reenact faces via boundary transfer
Wayne Wu, Yunxuan Zhang, Cheng Li, Chen Qian, and Chen Change Loy · 2018
Earlier work this paper cites.
Elegant: Exchanging latent encodings with gan for transferring multiple face attributes
Taihong Xiao, Jiapeng Hong, and Jinwen Ma · 2018
Earlier work this paper cites.
Arbitrary talking face generation via attentional audio-visual coherence learning
Hao Zhu, Huaibo Huang, Yi Li, Aihua Zheng, and Ran He · 2018
Earlier work this paper cites.
Hierarchical cross-modal talking face generation with dynamic pixel-wise loss
Lele Chen, Ross K Maddox, Zhiyao Duan, and Chenliang Xu · 2019
Earlier work this paper cites.
Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set
Yu Deng, Jiaolong Yang, Sicheng Xu, Dong Chen, Yunde Jia, and Xin Tong · 2019
Cited alongside, same era.
Infogan-cr: Disentangling generative adversarial networks with contrastive regularizers
Zinan Lin, Kiran Koshy Thekumparampil, Giulia C Fanti, and Sewoong Oh · 2019
Cited alongside, same era.
Challenging common assumptions in the unsupervised learning of disentangled representations
Francesco Locatello, Stefan Bauer, Mario Lucic, Gunnar Raetsch, Sylvain Gelly, Bernhard Schölkopf, and Olivier Bachem · 2019
Cited alongside, same era.
Hologan: Unsupervised learning of 3d representations from natural images
Thu Nguyen-Phuoc, Chuan Li, Lucas Theis, Christian Richardt, and Yong-Liang Yang · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Cited alongside, same era.
Live speech portraits: real-time photorealistic talking-head animation
Yuanxun Lu, Jinxiang Chai, and Xun Cao · 2021
Later among the works it cites.
Pirenderer: Controllable portrait image generation via semantic neural rendering
Yurui Ren, Ge Li, Yuanqi Chen, Thomas H Li, and Shan Liu · 2021
Later among the works it cites.
Closed-form factorization of latent semantics in gans
Yujun Shen and Bolei Zhou · 2021
Later among the works it cites.
One-shot free-view neural talking-head synthesis for video conferencing
Ting-Chun Wang, Arun Mallya, and Ming-Yu Liu · 2021
Later among the works it cites.
Orthogonal jacobian regularization for unsupervised disentanglement in image generation
Yuxiang Wei, Yupeng Shi, Xiao Liu, Zhilong Ji, Yuan Gao, Zhongqin Wu, and Wangmeng Zuo · 2021
Later among the works it cites.
Stylespace analysis: Disentangled controls for stylegan image generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
First order motion model for image animation
Aliaksandr Siarohin, Stéphane Lathuilière, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe · 2019
Cited alongside, same era.
Neural head reenactment with latent pose descriptors
Egor Burkov, Igor Pasechnik, Artur Grigorev, and Victor Lempitsky · 2020
Cited alongside, same era.
Talking-head generation with rhythmic head motion
Lele Chen, Guofeng Cui, Celong Liu, Zhong Li, Ziyi Kou, Yi Xu, and Chenliang Xu · 2020
Cited alongside, same era.
Disentangled and controllable face image generation via 3d imitative-contrastive learning
Yu Deng, Jiaolong Yang, Dong Chen, Fang Wen, and Xin Tong · 2020
Cited alongside, same era.
Guided variational autoencoder for disentanglement learning
Zheng Ding, Yifan Xu, Weijian Xu, Gaurav Parmar, Yang Yang, Max Welling, and Zhuowen Tu · 2020
Cited alongside, same era.
Gif: Generative interpretable faces
Partha Ghosh, Pravir Singh Gupta, Roy Uziel, Anurag Ranjan, Michael J Black, and Timo Bolkart · 2020
Cited alongside, same era.
Flnet: Landmark driven fetching and learning network for faithful talking facial animation synthesis
Kuangxiao Gu, Yuqian Zhou, and Thomas Huang · 2020
Cited alongside, same era.
Zongze Wu, Dani Lischinski, and Eli Shechtman · 2021
Later among the works it cites.
One-shot face reenactment using appearance adaptive normalization
Guangming Yao, Yi Yuan, Tianjia Shao, Shuang Li, Shanqi Liu, Yong Liu, Mengmeng Wang, and Kun Zhou · 2021
Later among the works it cites.
Heatmap regression via randomized rounding
Baosheng Yu and Dacheng Tao · 2021
Later among the works it cites.
Facial: Synthesizing dynamic talking face with implicit attribute learning
Chenxu Zhang, Yifan Zhao, Yifei Huang, Ming Zeng, Saifeng Ni, Madhukar Budagavi, and Xiaohu Guo · 2021
Later among the works it cites.
Real-time audio-guided multi-face reenactment
Jiangning Zhang, Xianfang Zeng, Chao Xu, Yong Liu, and Hongliang Li · 2021
Later among the works it cites.
Unsupervised disentanglement of linear-encoded facial semantics
Yutong Zheng, Yu-Kai Huang, Ran Tao, Zhiqiang Shen, and Marios Savvides · 2021
Later among the works it cites.
Pose-controllable talking face generation by implicitly modularized audio-visual representation
Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu · 2021
Later among the works it cites.
L2cs-net: Fine-grained gaze estimation in unconstrained environments
Ahmed A Abdelrahman, Thorsten Hempel, Aly Khalifa, and Ayoub Al-Hamadi · 2022
Closest in time.
Finding directions in gan’s latent space for neural face reenactment
Stella Bounareli, Vasileios Argyriou, and Georgios Tzimiropoulos · 2022
Closest in time.
Emoca: Emotion driven monocular face capture and animation
Radek Daněček, Michael J Black, and Timo Bolkart · 2022
Closest in time.
Megaportraits: One-shot megapixel neural head avatars
Nikita Drobyshev, Jenya Chelishev, Taras Khakhulin, Aleksei Ivakhnenko, Victor Lempitsky, and Egor Zakharov · 2022
Closest in time.
Depth-aware generative adversarial network for talking head video generation
Fa-Ting Hong, Longhao Zhang, Li Shen, and Dan Xu · 2022
Closest in time.
Eamm: One-shot emotional talking face via audio-based emotion-aware motion model
Xinya Ji, Hang Zhou, Kaisiyuan Wang, Qianyi Wu, Wayne Wu, Feng Xu, and Xun Cao · 2022
Closest in time.
Pushing the limits of raw waveform speaker recognition
Jee-weon Jung, You Jin Kim, Hee-Soo Heo, Bong-Jin Lee, Youngki Kwon, and Joon Son Chung · 2022
Closest in time.
Expressive talking head generation with granular audio-visual control
Borong Liang, Yan Pan, Zhizhi Guo, Hang Zhou, Zhibin Hong, Xiaoguang Han, Junyu Han, Jingtuo Liu, Errui Ding, and Jingdong Wang · 2022
Closest in time.
3d gan inversion for controllable portrait image animation
Connor Z Lin, David B Lindell, Eric R Chan, and Gordon Wetzstein · 2022
Closest in time.
Everybody’s talkin’: Let me talk as you want
Linsen Song, Wayne Wu, Chen Qian, Ran He, and Chen Change Loy · 2022
Closest in time.
Landmarkgan: Synthesizing faces from landmarks
Pu Sun, Yuezun Li, Honggang Qi, and Siwei Lyu · 2022
Closest in time.
Latent image animator: Learning to animate images via latent space navigation
Yaohui Wang, Di Yang, Francois Bremond, and Antitza Dantcheva · 2022
Closest in time.
Dfa-nerf: Personalized talking head generation via disentangled face attributes neural rendering
Shunyu Yao, RuiZhe Zhong, Yichao Yan, Guangtao Zhai, and Xiaokang Yang · 2022
Closest in time.
Styleheat: One-shot high-resolution editable talking face generation via pretrained stylegan
Fei Yin, Yong Zhang, Xiaodong Cun, Mingdeng Cao, Yanbo Fan, Xuan Wang, Qingyan Bai, Baoyuan Wu, Jue Wang, and Yujiu Yang · 2022
Closest in time.