Fetching the paper…
Reading the bibliography…
The existing methods for audio-driven talking head video editing have the limitations of poor visual effects.
I. Fodor, “Film dubbing: phonetic, semiotic, esthetic and psychological aspects,” 1976
1976
Earlier work this paper cites.
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,” IEEE transactions on image processing , vol. 13, no. 4, pp. 600–612, 2004
2004
Earlier work this paper cites.
JEREMY, SARACHAN, NANCI, BURK, KENNETH, DAY, MATTHEW, and TREVETT-SMITH, “Avatars talking: The use of virtual worlds within communication courses,” Journal of Interactive Learning Research , 2013
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
A. Bulat and G. Tzimiropoulos, “How far are we from solving the 2d & 3d face alignment problem?(and a dataset of 230,000 3d facial landmarks),” in International Conference on Computer Vision , 2017, pp. 1021–1030
2017
Earlier work this paper cites.
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 586–595
2018
Earlier work this paper cites.
T. Karras, S. Laine, and T. Aila, “A style-based generator architecture for generative adversarial networks,” in IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 4401–4410
2019
Earlier work this paper cites.
R. Abdal, Y. Qin, and P. Wonka, “Image2stylegan: How to embed images into the stylegan latent space?” in IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 4432–4441
2019
Earlier work this paper cites.
L. Chen, R. K. Maddox, Z. Duan, and C. Xu, “Hierarchical cross-modal talking face generation with dynamic pixel-wise loss,” in IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 7832–7841
2019
Earlier work this paper cites.
Y. Zhou, X. Han, E. Shechtman, J. Echevarria, E. Kalogerakis, and D. Li, “Makelttalk: speaker-aware talking-head animation,” ACM Transactions On Graphics (TOG) , vol. 39, no. 6, pp. 1–15, 2020
2020
Earlier work this paper cites.
E. Härkönen, A. Hertzmann, J. Lehtinen, and S. Paris, “Ganspace: Discovering interpretable gan controls,” Advances in neural information processing systems , vol. 33, pp. 9841–9850, 2020
2020
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial networks,” Communications of the ACM , vol. 63, no. 11, pp. 139–144, 2020
2020
Cited alongside, same era.
——, “Image2stylegan++: How to edit the embedded images?” in IEEE Conference on Computer Vision and Pattern Recognition , 2020, pp. 8296–8305
2020
Cited alongside, same era.
K. Wang, Q. Wu, L. Song, Z. Yang, W. Wu, C. Qian, R. He, Y. Qiao, and C. C. Loy, “Mead: A large-scale audio-visual dataset for emotional talking-face generation,” in European Conference on Computer Vision . Springer, 2020, pp. 700–717
2020
Cited alongside, same era.
K. Prajwal, R. Mukhopadhyay, V. P. Namboodiri, and C. Jawahar, “A lip sync expert is all you need for speech to lip generation in the wild,” in ACM International Conference on Multimedia , 2020, pp. 484–492
2020
Cited alongside, same era.
E. Richardson, Y. Alaluf, O. Patashnik, Y. Nitzan, Y. Azar, S. Shapiro, and D. Cohen-Or, “Encoding in style: a stylegan encoder for image-to-image translation,” in IEEE Conference on Computer Vision and Pattern Recognition , 2021, pp. 2287–2296
2021
Later among the works it cites.
O. Tov, Y. Alaluf, Y. Nitzan, O. Patashnik, and D. Cohen-Or, “Designing an encoder for stylegan image manipulation,” ACM Transactions on Graphics (TOG) , vol. 40, no. 4, pp. 1–14, 2021
2021
Later among the works it cites.
C. Yu, C. Gao, J. Wang, G. Yu, C. Shen, and N. Sang, “Bisenet v2: Bilateral network with guided aggregation for real-time semantic segmentation,” International Journal of Computer Vision , vol. 129, pp. 3051–3068, 2021
2021
Later among the works it cites.
Z. Zhang, L. Li, Y. Ding, and C. Fan, “Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset,” in IEEE Conference on Computer Vision and Pattern Recognition , 2021, pp. 3661–3670
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Lu, J. Chai, and X. Cao, “Live speech portraits: real-time photorealistic talking-head animation,” ACM Transactions on Graphics (TOG) , vol. 40, no. 6, pp. 1–17, 2021
2021
Cited alongside, same era.
X. Ji, H. Zhou, K. Wang, W. Wu, C. C. Loy, X. Cao, and F. Xu, “Audio-driven emotional video portraits,” in IEEE Conference on Computer Vision and Pattern Recognition , 2021, pp. 14 080–14 089
2021
Cited alongside, same era.
Y. Zhang, Y. Zhao, Y. Wen, Z. Tang, X. Xu, and M. Liu, “Facial prior based first order motion model for micro-expression generation,” in ACM International Conference on Multimedia , 2021, pp. 4755–4759
2021
Cited alongside, same era.
Y. Guo, K. Chen, S. Liang, Y.-J. Liu, H. Bao, and J. Zhang, “Ad-nerf: Audio driven neural radiance fields for talking head synthesis,” in IEEE Conference on Computer Vision and Pattern Recognition , 2021, pp. 5784–5794
2021
Cited alongside, same era.
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” Communications of the ACM , vol. 65, no. 1, pp. 99–106, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
R. Abdal, P. Zhu, N. J. Mitra, and P. Wonka, “Styleflow: Attribute-conditioned exploration of stylegan-generated images using conditional continuous normalizing flows,” ACM Transactions on Graphics (ToG) , vol. 40, no. 3, pp. 1–21, 2021
2021
Cited alongside, same era.
T. Xie, L. Liao, C. Bi, B. Tang, X. Yin, J. Yang, M. Wang, J. Yao, Y. Zhang, and Z. Ma, “Towards realistic visual dubbing with heterogeneous sources,” 2022
2022
Later among the works it cites.
F. Yin, Y. Zhang, X. Cun, M. Cao, Y. Fan, X. Wang, Q. Bai, B. Wu, J. Wang, and Y. Yang, “Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan,” in European Conference on Computer Vision . Springer, 2022, pp. 85–101
2022
Later among the works it cites.
R. Tzaban, R. Mokady, R. Gal, A. Bermano, and D. Cohen-Or, “Stitch it in time: Gan-based facial editing of real videos,” in SIGGRAPH , 2022, pp. 1–9
2022
Later among the works it cites.
2022
Later among the works it cites.
R. Gal, O. Patashnik, H. Maron, A. H. Bermano, G. Chechik, and D. Cohen-Or, “Stylegan-nada: Clip-guided domain adaptation of image generators,” ACM Transactions on Graphics (TOG) , vol. 41, no. 4, pp. 1–13, 2022
2022
Later among the works it cites.
D. Roich, R. Mokady, A. H. Bermano, and D. Cohen-Or, “Pivotal tuning for latent-based editing of real images,” ACM Transactions on graphics (TOG) , vol. 42, no. 1, pp. 1–13, 2022
2022
Later among the works it cites.
K. Cheng, X. Cun, Y. Zhang, M. Xia, F. Yin, M. Zhu, X. Wang, J. Wang, and N. Wang, “Videoretalking: Audio-based lip synchronization for talking head video editing in the wild,” in SIGGRAPH , 2022, pp. 1–9
2022
Later among the works it cites.