Fetching the paper…
Reading the bibliography…
We present a method to edit a target portrait footage by taking a sequence of audio as input to synthesize a photo-realistic video.
Speaker adaptation using constrained estimation of gaussian mixtures
Vassilios V. Digalakis, Dimitry Rtischev, and L. G. Neumeyer · 1995
Earlier work this paper cites.
Active appearance models
Timothy F. Cootes, Gareth J. Edwards, and Christopher J. Taylor · 1998
Earlier work this paper cites.
Maximum likelihood linear transformations for hmm-based speech recognition
Mark JF Gales · 1998
Earlier work this paper cites.
A morphable model for the synthesis of 3d faces
Volker Blanz and Thomas Vetter · 1999
Earlier work this paper cites.
Structuring linear transforms for adaptation using training time information
Karthik Visweswariah, Vaibhava Goel, and Ramesh Gopinath · 2002
Earlier work this paper cites.
Poisson image editing
Patrick Pérez, Michel Gangnet, and Andrew Blake · 2003
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, Eero P Simoncelli, et al · 2004
Earlier work this paper cites.
An audio-visual corpus for speech perception and automatic speech recognition
Martin Cooke, Stuart Cunningham, and Xu Shao · 2006
Earlier work this paper cites.
Visualizing data using t-sne
Laurens van der Maaten and Geoffrey E. Hinton · 2008
Earlier work this paper cites.
A 3d face model for pose and illumination invariant face recognition
Pascal Paysan, Reinhard Knothe, Brian Amberg, Sami Romdhani, and Thomas Vetter · 2009
Earlier work this paper cites.
The digital emily project: Achieving a photorealistic digital actor
Oleg Alexander, M. Rogers, W. Lambeth, Jen-Yuan Chiang, Wan-Chun Ma, Chuan-Chang Wang, and Paul E. Debevec · 2010
Earlier work this paper cites.
A basis representation of constrained mllr transforms for robust adaptation
Daniel Povey and Kaisheng Yao · 2012
Earlier work this paper cites.
Facewarehouse: A 3d facial expression database for visual computing
Chen Cao, Yanlin Weng, Shun Zhou, Yiying Tong, and Kun Zhou · 2014
Earlier work this paper cites.
Driving high-resolution facial scans with video performance capture
Graham Fyffe, Andrew Jones, Oleg Alexander, Ryosuke Ichikari, and Paul E. Debevec · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Blind video temporal consistency
Nicolas Bonneel, James Tompkin, Kalyan Sunkavalli, Deqing Sun, Sylvain Paris, and Hanspeter Pfister · 2015
Earlier work this paper cites.
Blind video temporal consistency
Nicolas Bonneel, James Tompkin, Kalyan Sunkavalli, Deqing Sun, Sylvain Paris, and Hanspeter Pfister · 2015
Earlier work this paper cites.
Vdub: Modifying face video of actors for plausible visual alignment to a dubbed audio track
Pablo Garrido, Levi Valgaerts, H. Sarmadi, Ingmar Steiner, Kiran Varanasi, Patrick Pérez, and Christian Theobalt · 2015
Earlier work this paper cites.
Unsupervised representation learning with deep convolutional generative adversarial networks
Alec Radford, Luke Metz, and Soumith Chintala · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Jali: an animator-centric viseme model for expressive lip synchronization
Peter Edwards, Chris Landreth, Eugene Fiume, and Karan Singh · 2016
Earlier work this paper cites.
Perceptual losses for real-time style transfer and super-resolution
Justin Johnson, Alexandre Alahi, and Li Fei-Fei · 2016
Cited alongside, same era.
Least squares generative adversarial networks
Xudong Mao, Qing Li, Haoran Xie, Raymond Y. K. Lau, Zhen Wang, and Stephen Paul Smolley · 2016
Cited alongside, same era.
Context encoders: Feature learning by inpainting
Deepak Pathak, Philipp Krähenbühl, Jeff Donahue, Trevor Darrell, and Alexei A. Efros · 2016
Cited alongside, same era.
Adaptive 3d face reconstruction from unconstrained photo collections
Joseph Roth, Yiying Tong, and Xiaoming Liu · 2016
Cited alongside, same era.
Audio-to-visual speech conversion using deep neural networks
Sarah Taylor, Akihiro Kato, Iain A. Matthews, and Ben P. Milner · 2016
Cited alongside, same era.
Face2face: Real-time face capture and reenactment of rgb videos
Justus Thies, Michael Zollhöfer, Marc Stamminger, Christian Theobalt, and Matthias Nießner · 2016
Conditional image generation for learning the structure of visual objects
Tomas Jakab, Ankush Gupta, Hakan Bilen, and Andrea Vedaldi · 2018
Later among the works it cites.
Deep video portraits
Hyeongwoo Kim, Pablo Garrido, Ayush Tewari, Weipeng Xu, Justus Thies, Matthias Nießner, Patrick Pérez, Christian Richardt, Michael Zollhöfer, and Christian Theobalt · 2018
Later among the works it cites.
Image inpainting for irregular holes using partial convolutions
Guilin Liu, Fitsum A. Reda, Kevin J. Shih, Ting-Chun Wang, Andrew Tao, and Bryan Catanzaro · 2018
Later among the works it cites.
pagan: real-time avatars using dynamic textures
Koki Nagano, Jaewoo Seo, Jun Xing, Lingyu Wei, Zimo Li, Shunsuke Saito, Aviral Agarwal, Jens Fursund, and Hao Li · 2018
Later among the works it cites.
Hai Xuan Pham, Yuting Wang, and Vladimir Pavlovic · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Generating videos with scene dynamics
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Cited alongside, same era.
Improved training of wasserstein gans
Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron C Courville · 2017
Cited alongside, same era.
Beyond face rotation: Global and local perception gan for photorealistic and identity preserving frontal view synthesis
Rui Huang, Shu Zhang, Tianyu Li, and Ran He · 2017
Cited alongside, same era.
Globally and locally consistent image completion
Satoshi Iizuka, Edgar Simo-Serra, and Hiroshi Ishikawa · 2017
Cited alongside, same era.
Audio-driven facial animation by joint end-to-end learning of pose and emotion
Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen · 2017
Cited alongside, same era.
Learning a model of facial shape and expression from 4d scans
Tianye Li, Timo Bolkart, Michael J. Black, Hao Li, and Javier Romero · 2017
Cited alongside, same era.
Geometry-aware face completion and editing
Linsen Song, Jie Cao, Linxiao Song, Yibo Hu, and Ran He · 2018
Later among the works it cites.
End-to-end speech-driven facial animation with temporal gans
Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic · 2018
Later among the works it cites.
X2face: A network for controlling face generation using images, audio, and pose codes
Olivia Wiles, A. Sophia Koepke, and Andrew Zisserman · 2018
Later among the works it cites.
Reenactgan: Learning to reenact faces via boundary transfer
Wayne Wu, Yunxuan Zhang, Cheng Li, Chen Qian, and Chen Change Loy · 2018
Later among the works it cites.
Free-form image inpainting with gated convolution
Jiahui Yu, Zhe L. Lin, Jimei Yang, Xiaohui Shen, Xin Lu, and Thomas S. Huang · 2018
Later among the works it cites.
Talking face generation by adversarially disentangled audio-visual representation
Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang · 2018
Later among the works it cites.
Visemenet: Audio-driven animator-centric speech animation
Yang Zhou, Zhan Xu, Chris Landreth, Evangelos Kalogerakis, Subhransu Maji, and Karan Singh · 2018
Later among the works it cites.
High-resolution talking face generation via mutual information approximation
Hao Zhu, Aihua Zheng, Huaibo Huang, and Ran He · 2018
Later among the works it cites.
Hierarchical cross-modal talking face generation with dynamic pixel-wise loss
Lele Chen, Ross K. Maddox, Zhiyao Duan, and Chenliang Xu · 2019
Later among the works it cites.
Text-based editing of talking-head video
Ohad Fried, Maneesh Agrawala, Ayush Tewari, Michael Zollhöfer, Adam Finkelstein, Eli Shechtman, Dan B Goldman, Kyle Genova, Zeyu Jin, and Christian Theobalt · 2019
Later among the works it cites.
You said that?: Synthesising talking faces from audio
Amir Jamaludin, Joon Son Chung, and Andrew Zisserman · 2019
Later among the works it cites.
Neural style-preserving visual dubbing
Hyeongwoo Kim, Mohamed Elgharib, Michael Zollhöfer, Hans-Peter Seidel, Thabo Beeler, Christian Richardt, and Christian Theobalt · 2019
Later among the works it cites.
Make a face: Towards arbitrary high fidelity face manipulation
Shengju Qian, Kwan-Yee Lin, Wayne Wu, Yangxiaokang Liu, Quan Wang, Fumin Shen, Chen Qian, and Ran He · 2019
Later among the works it cites.
Example-guided style consistent image synthesis from semantic labeling
Miao Wang, Guo-Ye Yang, Ruilong Li, Run-Ze Liang, Song-Hai Zhang, Peter. M. Hall, and Shi-Min Hu · 2019
Later among the works it cites.
Disentangling content and style via unsupervised geometry distillation
Wayne Wu, Kaidi Cao, Cheng Li, Chen Qian, and Chen Change Loy · 2019
Later among the works it cites.
Transgaga: Geometry-aware unsupervised image-to-image translation
Wayne Wu, Kaidi Cao, Cheng Li, Chen Qian, and Chen Change Loy · 2019
Later among the works it cites.