Fetching the paper…
Reading the bibliography…
One-shot talking face generation aims at synthesizing a high-quality talking face video from an arbitrary portrait image, driven by a video or an audio segment.
A morphable model for the synthesis of 3D faces. In SIGGRAPH
Volker Blanz and Thomas Vetter. 1999 · 1999
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
A 3d morphable model learnt from 10,000 faces. In CVPR
James Booth, Anastasios Roussos, Stefanos Zafeiriou, Allan Ponniah, and David Dunaway. 2016 · 2016
Earlier work this paper cites.
Out of time: automated lip sync in the wild. In ACCV
Joon Son Chung and Andrew Zisserman. 2016 · 2016
Earlier work this paper cites.
Perceptual losses for real-time style transfer and super-resolution. In ECCV
Justin Johnson, Alexandre Alahi, and Li Fei-Fei. 2016 · 2016
Earlier work this paper cites.
Generative visual manipulation on the natural image manifold. In ECCV
Jun-Yan Zhu, Philipp Krähenbühl, Eli Shechtman, and Alexei A Efros. 2016 · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017 · 2017
Earlier work this paper cites.
Arbitrary style transfer in real-time with adaptive instance normalization. In ICCV
Xun Huang and Serge Belongie. 2017 · 2017
Earlier work this paper cites.
Image-to-Image Translation with Conditional Adversarial Networks
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. 2017 · 2017
Earlier work this paper cites.
Voxceleb: a large-scale speaker identification dataset. In INTERSPEECH
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
Synthesizing obama: learning lip sync from audio
Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman. 2017 · 2017
Earlier work this paper cites.
A deep learning approach for generalized speech animation
Sarah Taylor, Taehwan Kim, Yisong Yue, Moshe Mahler, James Krahe, Anastasio Garcia Rodriguez, Jessica Hodgins, and Iain Matthews. 2017 · 2017
Earlier work this paper cites.
Recycle-gan: Unsupervised video retargeting. In ECCV
Aayush Bansal, Shugao Ma, Deva Ramanan, and Yaser Sheikh. 2018 · 2018
Earlier work this paper cites.
Warp-Guided GANs for Single-Photo Facial Animation
Jiahao Geng, Tianjia Shao, Youyi Zheng, Yanlin Weng, and Kun Zhou. 2018 · 2018
Earlier work this paper cites.
Progressive growing of gans for improved quality, stability, and variation. In ICLR
Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. 2018 · 2018
Earlier work this paper cites.
Deep video portraits
Hyeongwoo Kim, Pablo Garrido, Ayush Tewari, Weipeng Xu, Justus Thies, Matthias Niessner, Patrick Pérez, Christian Richardt, Michael Zollhöfer, and Christian Theobalt. 2018 · 2018
Earlier work this paper cites.
Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Guilin Liu, Andrew Tao, Jan Kautz, and Bryan Catanzaro. 2018a · 2018
Earlier work this paper cites.
X2face: A network for controlling face generation using images, audio, and pose codes. In ECCV
Olivia Wiles, A Koepke, and Andrew Zisserman. 2018 · 2018
Earlier work this paper cites.
Reenactgan: Learning to reenact faces via boundary transfer. In ECCV
Wayne Wu, Yunxuan Zhang, Cheng Li, Chen Qian, and Chen Change Loy. 2018 · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric. In CVPR
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. 2018 · 2018
Earlier work this paper cites.
Image2stylegan: How to embed images into the stylegan latent space?. In CVPR
Rameen Abdal, Yipeng Qin, and Peter Wonka. 2019 · 2019
Earlier work this paper cites.
Capture, learning, and synthesis of 3D speaking styles. In CVPR
Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael J Black. 2019 · 2019
Earlier work this paper cites.
Text-based editing of talking-head video
Ohad Fried, Ayush Tewari, Michael Zollhöfer, Adam Finkelstein, Eli Shechtman, Dan B Goldman, Kyle Genova, Zeyu Jin, Christian Theobalt, and Maneesh Agrawala. 2019 · 2019
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks. In CVPR
Tero Karras, Samuli Laine, and Timo Aila. 2019 · 2019
Cited alongside, same era.
Neural style-preserving visual dubbing
Hyeongwoo Kim, Mohamed Elgharib, Michael Zollhöfer, Hans-Peter Seidel, Thabo Beeler, Christian Richardt, and Christian Theobalt. 2019 · 2019
Cited alongside, same era.
First order motion model for image animation
Aliaksandr Siarohin, Stéphane Lathuilière, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. 2019a · 2019
Cited alongside, same era.
Few-shot video-to-video synthesis
Ting-Chun Wang, Ming-Yu Liu, Andrew Tao, Guilin Liu, Jan Kautz, and Bryan Catanzaro. 2019 · 2019
Cited alongside, same era.
Few-shot adversarial learning of realistic neural talking head models. In ICCV
Egor Zakharov, Aliaksandra Shysheya, Egor Burkov, and Victor Lempitsky. 2019 · 2019
Cited alongside, same era.
Stylevideogan: A temporal generative model using a pretrained stylegan
Gereon Fox, Ayush Tewari, Mohamed Elgharib, and Christian Theobalt. 2021 · 2021
Later among the works it cites.
AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis. In ICCV
Yudong Guo, Keyu Chen, Sen Liang, Yongjin Liu, Hujun Bao, and Juyong Zhang. 2021 · 2021
Later among the works it cites.
GAN Inversion for Out-of-Range Images with Geometric Transformations. In CVPR
Kyoungkook Kang, Seongtae Kim, and Sunghyun Cho. 2021 · 2021
Later among the works it cites.
Alias-free generative adversarial networks. In NIPS
Tero Karras, Miika Aittala, Samuli Laine, Erik Härkönen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2021 · 2021
Later among the works it cites.
Exploiting Spatial Dimensions of Latent in GAN for Real-Time Image Editing. In CVPR
Hyunsu Kim, Yunjey Choi, Junho Kim, Sungjoo Yoo, and Youngjung Uh. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Talking face generation by adversarially disentangled audio-visual representation. In AAAI
Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang. 2019 · 2019
Cited alongside, same era.
Image2StyleGAN++: How to Edit the Embedded Images?. In CVPR
Rameen Abdal, Yipeng Qin, and Peter Wonka. 2020 · 2020
Cited alongside, same era.
Neural head reenactment with latent pose descriptors. In CVPR
Egor Burkov, Igor Pasechnik, Artur Grigorev, and Victor Lempitsky. 2020 · 2020
Cited alongside, same era.
Sofgan: A portrait image generator with dynamic styling
Anpei Chen, Ruiyang Liu, Ling Xie, Zhang Chen, Hao Su, and Jingyi Yu. 2020 · 2020
Cited alongside, same era.
Analyzing and improving the image quality of stylegan. In CVPR
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2020 · 2020
Cited alongside, same era.
Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. 2020 · 2020
Cited alongside, same era.
A lip sync expert is all you need for speech to lip generation in the wild. In ACM Multimedia
KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar. 2020 · 2020
Cited alongside, same era.
LipSync3D: Data-Efficient Learning of Personalized 3D Talking Faces from Video using Pose and Lighting Normalization. In CVPR
Avisek Lahiri, Vivek Kwatra, Christian Frueh, John Lewis, and Chris Bregler. 2021 · 2021
Later among the works it cites.
Live Speech Portraits: Real-Time Photorealistic Talking-Head Animation
Yuanxun Lu, Jinxiang Chai, and Xun Cao. 2021 · 2021
Later among the works it cites.
PIRenderer: Controllable Portrait Image Generation via Semantic Neural Rendering. In ICCV
Yurui Ren, Ge Li, Yuanqi Chen, Thomas H Li, and Shan Liu. 2021 · 2021
Later among the works it cites.
Encoding in style: a stylegan encoder for image-to-image translation. In CVPR
Elad Richardson, Yuval Alaluf, Or Patashnik, Yotam Nitzan, Yaniv Azar, Stav Shapiro, and Daniel Cohen-Or. 2021 · 2021
Later among the works it cites.
Motion Representations for Articulated Animation. In CVPR
Aliaksandr Siarohin, Oliver J Woodford, Jian Ren, Menglei Chai, and Sergey Tulyakov. 2021 · 2021
Later among the works it cites.
StyleGAN-V: A Continuous Video Generator with the Price, Image Quality and Perks of StyleGAN2
Ivan Skorokhodov, Sergey Tulyakov, and Mohamed Elhoseiny. 2021 · 2021
Later among the works it cites.
AgileGAN: stylizing portraits by inversion-consistent transfer learning
Guoxian Song, Linjie Luo, Jing Liu, Wan-Chun Ma, Chunpong Lai, Chuanxia Zheng, and Tat-Jen Cham. 2021a · 2021
Later among the works it cites.
Everything’s Talkin’: Pareidolia Face Reenactment
Linsen Song, Wayne Wu, Chaoyou Fu, Chen Qian, Chen Change Loy, and Ran He. 2021b · 2021
Later among the works it cites.
A good image generator is what you need for high-resolution video synthesis. In ICLR
Yu Tian, Jian Ren, Menglei Chai, Kyle Olszewski, Xi Peng, Dimitris N Metaxas, and Sergey Tulyakov. 2021 · 2021
Later among the works it cites.
Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion
Suzhen Wang, Lincheng Li, Yu Ding, Changjie Fan, and Xin Yu. 2021b · 2021
Later among the works it cites.
One-shot Talking Face Generation from Single-speaker Audio-Visual Correlation Learning
Suzhen Wang, Lincheng Li, Yu Ding, and Xin Yu. 2021a · 2021
Later among the works it cites.
High-fidelity gan inversion for image attribute editing
Tengfei Wang, Yong Zhang, Yanbo Fan, Jue Wang, and Qifeng Chen. 2021e · 2021
Later among the works it cites.
A Simple Baseline for StyleGAN Inversion
Tianyi Wei, Dongdong Chen, Wenbo Zhou, Jing Liao, Weiming Zhang, Lu Yuan, Gang Hua, and Nenghai Yu. 2021 · 2021
Later among the works it cites.
Imitating Arbitrary Talking Style for Realistic Audio-Driven Talking Face Synthesis. In ACM Multimedia
Haozhe Wu, Jia Jia, Haoyu Wang, Yishun Dou, Chao Duan, and Qingshan Deng. 2021 · 2021
Later among the works it cites.
Pose-controllable talking face generation by implicitly modularized audio-visual representation. In CVPR
Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu. 2021 · 2021
Later among the works it cites.
Barbershop: GAN-based Image Compositing using Segmentation Masks
Peihao Zhu, Rameen Abdal, John Femiani, and Peter Wonka. 2021 · 2021
Later among the works it cites.
Latent Image Animator: Learning to animate image via latent space navigation. In ICLR
Anonymous. 2022 · 2022
Closest in time.
Stitch it in Time: GAN-Based Facial Editing of Real Videos
Rotem Tzaban, Ron Mokady, Rinon Gal, Amit H Bermano, and Daniel Cohen-Or. 2022 · 2022
Closest in time.