Fetching the paper…
Reading the bibliography…
Recently audio-driven talking face video generation has attracted considerable attention.
V. Blanz and T. Vetter, “A morphable model for the synthesis of 3d faces,” in SIGGRAPH , 1999, pp. 187–194
1999
Earlier work this paper cites.
P. Paysan, R. Knothe, B. Amberg, S. Romdhani, and T. Vetter, “A 3d face model for pose and illumination invariant face recognition,” in AVSS , 2009, pp. 296–301
2009
Earlier work this paper cites.
N. D. Narvekar and L. J. Karam, “A no-reference image blur metric based on the cumulative probability of blur detection (CPBD),” IEEE Trans. Image Process. , vol. 20, no. 9, pp. 2678–2683, 2011
2011
Earlier work this paper cites.
C. Cao, Y. Weng, S. Zhou, Y. Tong, and K. Zhou, “Facewarehouse: A 3d facial expression database for visual computing,” IEEE Trans. Vis. Comput. Graph. , vol. 20, no. 3, pp. 413–425, 2014
2014
Earlier work this paper cites.
P. Garrido, M. Zollhöfer, D. Casas, L. Valgaerts, K. Varanasi, P. Pérez, and C. Theobalt, “Reconstruction of personalized 3d face rigs from monocular video,” ACM Trans. Graph. , vol. 35, no. 3, pp. 28:1–28:15, 2016
2016
Earlier work this paper cites.
L. Theis, A. van den Oord, and M. Bethge, “A note on the evaluation of generative models,” in Proc. Int. Conf. Learn. Representations , 2016, pp. 1–10
2016
Earlier work this paper cites.
J. S. Chung and A. Zisserman, “Out of time: Automated lip sync in the wild,” in ACCV Workshops (2) , vol. 10117, 2016, pp. 251–263
2016
Earlier work this paper cites.
P. Isola, J. Zhu, T. Zhou, and A. A. Efros, “Image-to-image translation with conditional adversarial networks,” in CVPR , 2017, pp. 5967–5976
2017
Earlier work this paper cites.
J. Zhu, T. Park, P. Isola, and A. A. Efros, “Unpaired image-to-image translation using cycle-consistent adversarial networks,” in ICCV , 2017, pp. 2242–2251
2017
Earlier work this paper cites.
T. Li, T. Bolkart, M. J. Black, H. Li, and J. Romero, “Learning a model of facial shape and expression from 4d scans,” ACM Trans. Graph. , vol. 36, no. 6, pp. 194:1–194:17, 2017
2017
Earlier work this paper cites.
Z. C. Lipton and S. Tripathi, “Precise recovery of latent vectors from generative adversarial networks,” in ICLR (Workshop) . OpenReview.net, 2017
2017
Earlier work this paper cites.
V. Dumoulin, I. Belghazi, B. Poole, A. Lamb, M. Arjovsky, O. Mastropietro, and A. C. Courville, “Adversarially learned inference,” in Proc. Int. Conf. Learn. Representations , 2017, pp. 1–13
2017
Earlier work this paper cites.
T. Karras, T. Aila, S. Laine, A. Herva, and J. Lehtinen, “Audio-driven facial animation by joint end-to-end learning of pose and emotion,” ACM Trans. Graph. , vol. 36, no. 4, pp. 94:1–94:12, 2017
2017
Earlier work this paper cites.
M. Arjovsky, S. Chintala, and L. Bottou, “Wasserstein generative adversarial networks,” in ICML , vol. 70, 2017, pp. 214–223
2017
Earlier work this paper cites.
I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville, “Improved training of wasserstein gans,” pp. 5767–5777, 2017
2017
Earlier work this paper cites.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” pp. 6626–6637, 2017
2017
Earlier work this paper cites.
T. Li, T. Bolkart, M. J. Black, H. Li, and J. Romero, “Learning a model of facial shape and expression from 4d scans,” ACM Trans. Graph. , vol. 36, no. 6, pp. 194:1–194:17, 2017
2017
Earlier work this paper cites.
Y. Choi, M. Choi, M. Kim, J. Ha, S. Kim, and J. Choo, “Stargan: Unified generative adversarial networks for multi-domain image-to-image translation,” in CVPR , 2018, pp. 8789–8797
2018
Earlier work this paper cites.
L. Jiang, J. Zhang, B. Deng, H. Li, and L. Liu, “3d face reconstruction with geometry details from a single image,” IEEE Trans. Image Process. , vol. 27, no. 10, pp. 4756–4770, 2018
2018
Earlier work this paper cites.
H. Ding, K. Sricharan, and R. Chellappa, “Exprgan: Facial expression editing with controllable expression intensity,” in AAAI , 2018, pp. 6781–6788
2018
Earlier work this paper cites.
S. R. Livingstone and F. A. Russo, “The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,” PLOS ONE , vol. 13, no. 5, pp. 1–35, 05 2018
2018
Earlier work this paper cites.
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in CVPR , 2018, pp. 586–595
2018
Cited alongside, same era.
L. Chen, R. K. Maddox, Z. Duan, and C. Xu, “Hierarchical cross-modal talking face generation with dynamic pixel-wise loss,” in CVPR , 2019, pp. 7832–7841
2019
Cited alongside, same era.
T. Karras, S. Laine, and T. Aila, “A style-based generator architecture for generative adversarial networks,” in CVPR , 2019, pp. 4401–4410
2019
Cited alongside, same era.
Y. Deng, J. Yang, S. Xu, D. Chen, Y. Jia, and X. Tong, “Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set,” in CVPR Workshops , 2019, pp. 285–295
2019
Cited alongside, same era.
Z. Geng, C. Cao, and S. Tulyakov, “3d guided fine-grained face manipulation,” in CVPR , 2019, pp. 9821–9830
K. R. Prajwal, R. Mukhopadhyay, V. P. Namboodiri, and C. V. Jawahar, “A lip sync expert is all you need for speech to lip generation in the wild,” in ACM Multimedia , 2020, pp. 484–492
2020
Later among the works it cites.
A. Lahiri, V. Kwatra, C. Früh, J. Lewis, and C. Bregler, “Lipsync3d: Data-efficient learning of personalized 3d talking faces from video using pose and lighting normalization,” in CVPR , 2021, pp. 2755–2764
2021
Later among the works it cites.
X. Ji, H. Zhou, K. Wang, W. Wu, C. C. Loy, X. Cao, and F. Xu, “Audio-driven emotional video portraits,” in CVPR , 2021, pp. 14 080–14 089
2021
Later among the works it cites.
E. Richardson, Y. Alaluf, O. Patashnik, Y. Nitzan, Y. Azar, S. Shapiro, and D. Cohen-Or, “Encoding in style: A stylegan encoder for image-to-image translation,” in CVPR , 2021, pp. 2287–2296
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
R. Abdal, Y. Qin, and P. Wonka, “Image2stylegan: How to embed images into the stylegan latent space?” in ICCV , 2019, pp. 4431–4440
2019
Cited alongside, same era.
L. Ma and Z. Deng, “Real-time facial expression transformation for monocular RGB video,” Comput. Graph. Forum , vol. 38, no. 1, pp. 470–481, 2019
2019
Cited alongside, same era.
J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “Arcface: Additive angular margin loss for deep face recognition,” in CVPR , 2019, pp. 4690–4699
2019
Cited alongside, same era.
A. Mollahosseini, B. Hassani, and M. H. Mahoor, “Affectnet: A database for facial expression, valence, and arousal computing in the wild,” IEEE Trans. Affect. Comput. , vol. 10, no. 1, pp. 18–31, 2019
2019
Cited alongside, same era.
J. Thies, M. Elgharib, A. Tewari, C. Theobalt, and M. Nießner, “Neural voice puppetry: Audio-driven facial reenactment,” in ECCV (16) , vol. 12361, 2020, pp. 716–731
2020
Cited alongside, same era.
K. Vougioukas, S. Petridis, and M. Pantic, “Realistic speech-driven facial animation with gans,” Int. J. Comput. Vis. , vol. 128, no. 5, pp. 1398–1413, 2020
2020
Cited alongside, same era.
K. R. Prajwal, R. Mukhopadhyay, V. P. Namboodiri, and C. V. Jawahar, “A lip sync expert is all you need for speech to lip generation in the wild,” in ACM Multimedia , 2020, pp. 484–492
2020
Cited alongside, same era.
O. Tov, Y. Alaluf, Y. Nitzan, O. Patashnik, and D. Cohen-Or, “Designing an encoder for stylegan image manipulation,” ACM Trans. Graph. , vol. 40, no. 4, pp. 133:1–133:14, 2021
2021
Later among the works it cites.
Y. Alaluf, O. Patashnik, and D. Cohen-Or, “Restyle: A residual-based stylegan encoder via iterative refinement,” in ICCV , 2021, pp. 6691–6700
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. Feng, H. Feng, M. J. Black, and T. Bolkart, “Learning an animatable detailed 3d face model from in-the-wild images,” ACM Trans. Graph. , vol. 40, no. 4, pp. 88:1–88:13, 2021
2021
Later among the works it cites.
S. d’Apolito, D. P. Paudel, Z. Huang, A. Romero, and L. V. Gool, “Ganmut: Learning interpretable conditional space for gamut of emotions,” in CVPR . Computer Vision Foundation / IEEE, 2021, pp. 568–577
2021
Later among the works it cites.
Y. Shen and B. Zhou, “Closed-form factorization of latent semantics in gans,” in CVPR , 2021, pp. 1532–1540
2021
Later among the works it cites.
I. Magnusson, A. Sankaranarayanan, and A. Lippman, “Invertable frowns: Video-to-video facial emotion translation,” in ADGD @ ACM Multimedia , 2021, pp. 25–33
2021
Later among the works it cites.
T. Karras, M. Aittala, S. Laine, E. Härkönen, J. Hellsten, J. Lehtinen, and T. Aila, “Alias-free generative adversarial networks,” in Proc. NeurIPS , 2021, pp. 852–863
2021
Later among the works it cites.
A. V. Savchenko, “Facial expression and attributes recognition based on multi-task learning of lightweight neural networks,” in SISY , 2021, pp. 119–124
2021
Later among the works it cites.
Y. Feng, H. Feng, M. J. Black, and T. Bolkart, “Learning an animatable detailed 3d face model from in-the-wild images,” ACM Trans. Graph. , vol. 40, no. 4, pp. 88:1–88:13, 2021
2021
Later among the works it cites.
R. Yi, Z. Ye, Z. Sun, J. Zhang, G. Zhang, P. Wan, H. Bao, and Y.-J. Liu, “Predicting personalized head movement from short video and speech signal,” IEEE Trans. Multimedia , pp. 1–13, 2022
2022
Closest in time.
2022
Closest in time.
F. P. Papantoniou, P. P. Filntisis, P. Maragos, and A. Roussos, “Neural emotion director: Speech-preserving semantic control of facial expressions in ”in-the-wild” videos,” in CVPR . IEEE, 2022, pp. 18 759–18 768
2022
Closest in time.
T. Wang, Y. Zhang, Y. Fan, J. Wang, and Q. Chen, “High-fidelity GAN inversion for image attribute editing,” in CVPR . IEEE, 2022, pp. 11 369–11 378
2022
Closest in time.
G. K. Solanki and A. Roussos, “Deep semantic manipulation of facial videos,” in ECCV Workshops (6) , ser. Lecture Notes in Computer Science, vol. 13806. Springer, 2022, pp. 104–120
2022
Closest in time.
F.-L. Liu, S.-Y. Chen, Y.-K. Lai, C. Li, Y.-R. Jiang, H. Fu, and L. Gao, “DeepFaceVideoEditing: Sketch-based deep editing of face videos,” ACM Transactions on Graphics , vol. 41, no. 4, pp. 167:1–167:16, 2022
2022
Closest in time.