Fetching the paper…
Reading the bibliography…
For few-shot learning, it is still a critical challenge to realize photo-realistic face visually dubbing on high-resolution videos.
Image quality assessment: from error visibility to structural similarity
Wang, Z.; Bovik, A. C.; Sheikh, H. R.; and Simoncelli, E. P. 2004 · 2004
Earlier work this paper cites.
Deep speech: Scaling up end-to-end speech recognition
Hannun, A.; Case, C.; Casper, J.; Catanzaro, B.; Diamos, G.; Elsen, E.; Prenger, R.; Satheesh, S.; Sengupta, S.; Coates, A.; et al. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K.; and Zisserman, A. 2014 · 2014
Earlier work this paper cites.
Spatial transformer networks
Jaderberg, M.; Simonyan, K.; Zisserman, A.; et al. 2015 · 2015
Earlier work this paper cites.
Openface: an open source facial behavior analysis toolkit
Baltrušaitis, T.; Robinson, P.; and Morency, L.-P. 2016 · 2016
Earlier work this paper cites.
Perceptual losses for real-time style transfer and super-resolution
Johnson, J.; Alahi, A.; and Fei-Fei, L. 2016 · 2016
Earlier work this paper cites.
Chung, J. S.; Jamaludin, A.; and Zisserman, A. 2017 · 2017
Earlier work this paper cites.
Arbitrary style transfer in real-time with adaptive instance normalization
Huang, X.; and Belongie, S. 2017 · 2017
Earlier work this paper cites.
Inverse compositional spatial transformer networks
Lin, C.-H.; and Lucey, S. 2017 · 2017
Earlier work this paper cites.
Least squares generative adversarial networks
Mao, X.; Li, Q.; Xie, H.; Lau, R. Y.; Wang, Z.; and Paul Smolley, S. 2017 · 2017
Earlier work this paper cites.
Synthesizing obama: learning lip sync from audio
Suwajanakorn, S.; Seitz, S. M.; and Kemelmacher-Shlizerman, I. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Lip movements generation at a glance
Chen, L.; Li, Z.; Maddox, R. K.; Duan, Z.; and Xu, C. 2018 · 2018
Earlier work this paper cites.
Talking face generation by conditional recurrent adversarial network
Song, Y.; Zhu, J.; Li, D.; Wang, X.; and Qi, H. 2018 · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R.; Isola, P.; Efros, A. A.; Shechtman, E.; and Wang, O. 2018 · 2018
Earlier work this paper cites.
Hierarchical cross-modal talking face generation with dynamic pixel-wise loss
Chen, L.; Maddox, R. K.; Duan, Z.; and Xu, C. 2019 · 2019
Cited alongside, same era.
Text-based editing of talking-head video
Fried, O.; Tewari, A.; Zollhöfer, M.; Finkelstein, A.; Shechtman, E.; Goldman, D. B.; Genova, K.; Jin, Z.; Theobalt, C.; and Agrawala, M. 2019 · 2019
Cited alongside, same era.
Towards automatic face-to-face translation
KR, P.; Mukhopadhyay, R.; Philip, J.; Jha, A.; Namboodiri, V.; and Jawahar, C. 2019 · 2019
Cited alongside, same era.
First order motion model for image animation
Siarohin, A.; Lathuilière, S.; Tulyakov, S.; Ricci, E.; and Sebe, N. 2019 · 2019
Cited alongside, same era.
Talking face generation by adversarially disentangled audio-visual representation
Zhou, H.; Liu, Y.; Liu, Z.; Luo, P.; and Wang, X. 2019 · 2019
Cited alongside, same era.
Talking-head generation with rhythmic head motion
Chen, L.; Cui, G.; Liu, C.; Li, Z.; Kou, Z.; Xu, Y.; and Xu, C. 2020 · 2020
Ad-nerf: Audio driven neural radiance fields for talking head synthesis
Guo, Y.; Chen, K.; Liang, S.; Liu, Y.-J.; Bao, H.; and Zhang, J. 2021 · 2021
Later among the works it cites.
Audio-driven emotional video portraits
Ji, X.; Zhou, H.; Wang, K.; Wu, W.; Loy, C. C.; Cao, X.; and Xu, F. 2021 · 2021
Later among the works it cites.
Lipsync3d: Data-efficient learning of personalized 3d talking faces from video using pose and lighting normalization
Lahiri, A.; Kwatra, V.; Frueh, C.; Lewis, J.; and Bregler, C. 2021 · 2021
Later among the works it cites.
Pirenderer: Controllable portrait image generation via semantic neural rendering
Ren, Y.; Li, G.; Chen, Y.; Li, T. H.; and Liu, S. 2021 · 2021
Later among the works it cites.
Audio2head: Audio-driven one-shot talking-head generation with natural head motion
Wang, S.; Li, L.; Ding, Y.; Fan, C.; and Yu, X. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Speech-driven facial animation using cascaded gans for learning of motion and texture
Das, D.; Biswas, S.; Sinha, S.; and Bhowmick, B. 2020 · 2020
Cited alongside, same era.
Nerf: Representing scenes as neural radiance fields for view synthesis
Mildenhall, B.; Srinivasan, P. P.; Tancik, M.; Barron, J. T.; Ramamoorthi, R.; and Ng, R. 2020 · 2020
Cited alongside, same era.
A lip sync expert is all you need for speech to lip generation in the wild
Prajwal, K.; Mukhopadhyay, R.; Namboodiri, V. P.; and Jawahar, C. 2020 · 2020
Cited alongside, same era.
Neural voice puppetry: Audio-driven facial reenactment
Thies, J.; Elgharib, M.; Tewari, A.; Theobalt, C.; and Nießner, M. 2020 · 2020
Cited alongside, same era.
Realistic speech-driven facial animation with gans
Vougioukas, K.; Petridis, S.; and Pantic, M. 2020 · 2020
Cited alongside, same era.
MEAD: A Large-scale Audio-visual Dataset for Emotional Talking-face Generation
Wang, K.; Wu, Q.; Song, L.; Yang, Z.; Wu, W.; Qian, C.; He, R.; Qiao, Y.; and Loy, C. C. 2020 · 2020
Cited alongside, same era.
One-shot free-view neural talking-head synthesis for video conferencing
Wang, T.-C.; Mallya, A.; and Liu, M.-Y. 2021 · 2021
Later among the works it cites.
Imitating arbitrary talking style for realistic audio-driven talking face synthesis
Wu, H.; Jia, J.; Wang, H.; Dou, Y.; Duan, C.; and Deng, Q. 2021 · 2021
Later among the works it cites.
Towards realistic visual dubbing with heterogeneous sources
Xie, T.; Liao, L.; Bi, C.; Tang, B.; Yin, X.; Yang, J.; Wang, M.; Yao, J.; Zhang, Y.; and Ma, Z. 2021 · 2021
Later among the works it cites.
Pose-controllable talking face generation by implicitly modularized audio-visual representation
Zhou, H.; Sun, Y.; Wu, W.; Loy, C. C.; Wang, X.; and Liu, Z. 2021 · 2021
Later among the works it cites.
EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion Model
Ji, X.; Zhou, H.; Wang, K.; Wu, Q.; Wu, W.; Xu, F.; and Cao, X. 2022 · 2022
Later among the works it cites.
Expressive talking head generation with granular audio-visual control
Liang, B.; Pan, Y.; Guo, Z.; Zhou, H.; Hong, Z.; Han, X.; Han, J.; Liu, J.; Ding, E.; and Wang, J. 2022 · 2022
Later among the works it cites.
Parallel and High-Fidelity Text-to-Lip Generation
Liu, J.; Zhu, Z.; Ren, Y.; Huang, W.; Huai, B.; Yuan, N.; and Zhao, Z. 2022 · 2022
Later among the works it cites.
SyncTalkFace: Talking Face Generation with Precise Lip-syncing via Audio-Lip Memory
Park, S. J.; Kim, M.; Hong, J.; Choi, J.; and Ro, Y. M. 2022 · 2022
Later among the works it cites.
Everybody’s talkin’: Let me talk as you want
Song, L.; Wu, W.; Qian, C.; He, R.; and Loy, C. C. 2022 · 2022
Later among the works it cites.
Adaptive Affine Transformation: A Simple and Effective Operation for Spatial Misaligned Image Generation
Zhang, Z.; and Ding, Y. 2022 · 2022
Later among the works it cites.