Fetching the paper…
Reading the bibliography…
The goal of this paper is to synthesise talking faces with controllable facial motions.
Multiscale structural similarity for image quality assessment. In The Asilomar Conference on Signals, Systems & Computers , Vol. 2. Ieee, 1398–1402
Zhou Wang, Eero P Simoncelli, and Alan C Bovik. 2003 · 2003
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. 2004 · 2004
Earlier work this paper cites.
An expressive text-driven 3d talking head
Robert Anderson, Björn Stenger, Vincent Wan, and Roberto Cipolla. 2013 · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Unsupervised discovery of interpretable directions in the gan latent space
Andrey Voynov and Artem Babenko. 2014 · 2014
Earlier work this paper cites.
Photo-real talking head with deep bidirectional LSTM. In Proc. ICASSP . IEEE, 4884–4888
Bo Fan, Lijuan Wang, Frank K Soong, and Lei Xie. 2015 · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition. In Proc. ICLR
Karen Simonyan and Andrew Zisserman. 2015 · 2015
Earlier work this paper cites.
A deep bidirectional LSTM approach for video-realistic talking head
Bo Fan, Lei Xie, Shan Yang, Lijuan Wang, and Frank K Soong. 2016 · 2016
Earlier work this paper cites.
How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks). In Proc. ICCV . 1021–1030
Adrian Bulat and Georgios Tzimiropoulos. 2017 · 2017
Earlier work this paper cites.
You said that?. In Proc. BMVC
Joon Son Chung, Amir Jamaludin, and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
Out of time: automated lip sync in the wild. In Proc. ACCV . Springer, 251–263
Joon Son Chung and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
Synthesizing obama: learning lip sync from audio
Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman. 2017 · 2017
Earlier work this paper cites.
Lip movements generation at a glance. In Proc. ECCV . 520–535
Lele Chen, Zhiheng Li, Ross K Maddox, Zhiyao Duan, and Chenliang Xu. 2018 · 2018
Earlier work this paper cites.
VoxCeleb2: Deep speaker recognition. In Proc. Interspeech
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman. 2018 · 2018
Earlier work this paper cites.
Talking face generation by conditional recurrent adversarial network
Yang Song, Jingwen Zhu, Dawei Li, Xiaolong Wang, and Hairong Qi. 2018 · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric. In Proc. CVPR . 586–595
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. 2018 · 2018
Earlier work this paper cites.
Hierarchical cross-modal talking face generation with dynamic pixel-wise loss. In Proc. CVPR . 7832–7841
Lele Chen, Ross K Maddox, Zhiyao Duan, and Chenliang Xu. 2019 · 2019
Earlier work this paper cites.
Ganalyze: Toward visual definitions of cognitive image properties. In Proc. ICCV . 5744–5753
Lore Goetschalckx, Alex Andonian, Aude Oliva, and Phillip Isola. 2019 · 2019
Earlier work this paper cites.
On the" steerability" of generative adversarial networks
Ali Jahanian, Lucy Chai, and Phillip Isola. 2019 · 2019
Earlier work this paper cites.
Disentangled representation learning for 3D face shape. In Proc. CVPR . 11957–11966
Zi-Hang Jiang, Qianyi Wu, Keyu Chen, and Juyong Zhang. 2019 · 2019
Cited alongside, same era.
Towards automatic face-to-face translation. In Proc. ACM MM . 1428–1436
Prajwal KR, Rudrabha Mukhopadhyay, Jerin Philip, Abhishek Jha, Vinay Namboodiri, and CV Jawahar. 2019 · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Cited alongside, same era.
First order motion model for image animation. In NeurIPS , Vol. 32
Aliaksandr Siarohin, Stéphane Lathuilière, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. 2019 · 2019
Cited alongside, same era.
Talking face generation by adversarially disentangled audio-visual representation. In Proc. AAAI , Vol. 33. 9299–9306
Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang. 2019 · 2019
Cited alongside, same era.
Arbitrary talking face generation via attentional audio-visual coherence learning. In Proc. IJCAI
Hao Zhu, Huaibo Huang, Yi Li, Aihua Zheng, and Ran He. 2020 · 2020
Later among the works it cites.
Audio-driven emotional video portraits. In Proc. CVPR . 14080–14089
Xinya Ji, Hang Zhou, Kaisiyuan Wang, Wayne Wu, Chen Change Loy, Xun Cao, and Feng Xu. 2021 · 2021
Later among the works it cites.
Audio-and gaze-driven facial animation of codec avatars. In Proc. WACV . 41–50
Alexander Richard, Colin Lea, Shugao Ma, Jurgen Gall, Fernando De la Torre, and Yaser Sheikh. 2021 · 2021
Later among the works it cites.
Encoding in style: a stylegan encoder for image-to-image translation. In Proc. CVPR . 2287–2296
Elad Richardson, Yuval Alaluf, Or Patashnik, Yotam Nitzan, Yaniv Azar, Stav Shapiro, and Daniel Cohen-Or. 2021 · 2021
Later among the works it cites.
Closed-form factorization of latent semantics in gans. In Proc. CVPR . 1532–1540
Yujun Shen and Bolei Zhou. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neural head reenactment with latent pose descriptors. In Proc. CVPR . 13786–13795
Egor Burkov, Igor Pasechnik, Artur Grigorev, and Victor Lempitsky. 2020 · 2020
Cited alongside, same era.
What comprises a good talking-head video generation?: A survey and benchmark
Lele Chen, Guofeng Cui, Ziyi Kou, Haitian Zheng, and Chenliang Xu. 2020a · 2020
Cited alongside, same era.
Speech-driven facial animation using cascaded gans for learning of motion and texture. In Proc. ECCV . Springer, 408–424
Dipanjan Das, Sandika Biswas, Sanjana Sinha, and Brojeshwar Bhowmick. 2020 · 2020
Cited alongside, same era.
Generative adversarial networks
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2020 · 2020
Cited alongside, same era.
On the ”steerability” of generative adversarial networks. In International Conference on Learning Representations
Ali Jahanian, Lucy Chai, and Phillip Isola. 2020 · 2020
Cited alongside, same era.
Analyzing and improving the image quality of stylegan. In Proc. CVPR . 8110–8119
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2020 · 2020
Cited alongside, same era.
The hessian penalty: A weak prior for unsupervised disentanglement. In Proc. ECCV . Springer, 581–597
William Peebles, John Peebles, Jun-Yan Zhu, Alexei Efros, and Antonio Torralba. 2020 · 2020
Cited alongside, same era.
Motion representations for articulated animation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 13653–13662
Aliaksandr Siarohin, Oliver J Woodford, Jian Ren, Menglei Chai, and Sergey Tulyakov. 2021 · 2021
Later among the works it cites.
Designing an encoder for stylegan image manipulation
Omer Tov, Yuval Alaluf, Yotam Nitzan, Or Patashnik, and Daniel Cohen-Or. 2021 · 2021
Later among the works it cites.
L2m-gan: Learning to manipulate latent space semantics for facial attribute editing. In Proc. CVPR . 2951–2960
Guoxing Yang, Nanyi Fei, Mingyu Ding, Guangzhen Liu, Zhiwu Lu, and Tao Xiang. 2021 · 2021
Later among the works it cites.
Pose-controllable talking face generation by implicitly modularized audio-visual representation. In Proc. CVPR . 4176–4186
Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu. 2021 · 2021
Later among the works it cites.
Deep audio-visual learning: A survey
Hao Zhu, Man-Di Luo, Rui Wang, Ai-Hua Zheng, and Ran He. 2021 · 2021
Later among the works it cites.
Hyperstyle: Stylegan inversion with hypernetworks for real image editing. In Proc. CVPR . 18511–18521
Yuval Alaluf, Omer Tov, Ron Mokady, Rinon Gal, and Amit Bermano. 2022 · 2022
Later among the works it cites.
Expressive talking head generation with granular audio-visual control. In Proc. CVPR . 3387–3396
Borong Liang, Yan Pan, Zhizhi Guo, Hang Zhou, Zhibin Hong, Xiaoguang Han, Junyu Han, Jingtuo Liu, Errui Ding, and Jingdong Wang. 2022 · 2022
Later among the works it cites.
StyleTalker: One-shot Style-based Audio-driven Talking Head Video Generation
Dongchan Min, Minyoung Song, and Sung Ju Hwang. 2022 · 2022
Later among the works it cites.
Everybody’s talkin’: Let me talk as you want
Linsen Song, Wayne Wu, Chen Qian, Ran He, and Chen Change Loy. 2022 · 2022
Later among the works it cites.
Duomin Wang, Yu Deng, Zixin Yin, Heung-Yeung Shum, and Baoyuan Wang. 2022a · 2022
Later among the works it cites.
Geumbyeol Hwang, Sunwon Hong, Seunghyun Lee, Sungwoo Park, and Gyeongsu Chae. 2023 · 2023
Closest in time.
StyleTalk: One-shot Talking Head Generation with Controllable Speaking Styles
Yifeng Ma, Suzhen Wang, Zhipeng Hu, Changjie Fan, Tangjie Lv, Yu Ding, Zhidong Deng, and Xin Yu. 2023 · 2023
Closest in time.
Synctalkface: Talking face generation with precise lip-syncing via audio-lip memory. In Proc. AAAI , Vol. 36. 2062–2070
Se Jin Park, Minsu Kim, Joanna Hong, Jeongsoo Choi, and Yong Man Ro. 2022 · 2070
Closest in time.