Fetching the paper…
Reading the bibliography…
While previous speech-driven talking face generation methods have made significant progress in improving the visual quality and lip-sync quality of the synthesized videos, they pay less attention to lip motion jitters which greatly undermine the realness of talking face videos.
W. Chen, X. Tan, Y. Xia, T. Qin, Y. Wang, and T.-Y. Liu, “Duallip: A system for joint lip reading and generation,” in Proc. ACM MM , 2020, pp. 1985–1993
1993
Earlier work this paper cites.
V. Blanz and T. Vetter, “A morphable model for the synthesis of 3d faces,” in Proceedings of the 26th annual conference on Computer graphics and interactive techniques , 1999, pp. 187–194
1999
Earlier work this paper cites.
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,” TIP , vol. 13, no. 4, pp. 600–612, 2004
2004
Earlier work this paper cites.
Y. Matsushita, E. Ofek, W. Ge, X. Tang, and H.-Y. Shum, “Full-frame video stabilization with motion inpainting,” TPAMI , vol. 28, no. 7, pp. 1150–1163, 2006
2006
Earlier work this paper cites.
D. Fleet and Y. Weiss, “Optical flow estimation,” in Handbook of mathematical models in computer vision . Springer, 2006, pp. 237–257
2006
Earlier work this paper cites.
M. Cooke, J. Barker, S. Cunningham, and X. Shao, “An audio-visual corpus for speech perception and automatic speech recognition,” The Journal of the Acoustical Society of America , vol. 120, no. 5, pp. 2421–2424, 2006
2006
Earlier work this paper cites.
K. Seshadrinathan and A. C. Bovik, “Motion-based perceptual quality assessment of video,” in Human Vision and Electronic Imaging XIV , vol. 7240. International Society for Optics and Photonics, 2009, p. 72400X
2009
Earlier work this paper cites.
F. Liu, M. Gleicher, H. Jin, and A. Agarwala, “Content-preserving warps for 3d video stabilization,” ACM TOG , vol. 28, no. 3, pp. 1–9, 2009
2009
Earlier work this paper cites.
N. D. Narvekar and L. J. Karam, “A no-reference perceptual image sharpness metric based on a cumulative probability of blur detection,” in 2009 International Workshop on Quality of Multimedia Experience . IEEE, 2009, pp. 87–91
2009
Earlier work this paper cites.
J. Thies, M. Zollhofer, M. Stamminger, C. Theobalt, and M. Nießner, “Face2face: Real-time face capture and reenactment of rgb videos,” in Proc. CVPR , 2016, pp. 2387–2395
2016
Earlier work this paper cites.
D. Amodei, S. Ananthanarayanan, R. Anubhai, J. Bai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, Q. Cheng, G. Chen et al. , “Deep speech 2: End-to-end speech recognition in english and mandarin,” in Proc. ICML . PMLR, 2016, pp. 173–182
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. CVPR , 2016, pp. 770–778
2016
Earlier work this paper cites.
J. Johnson, A. Alahi, and L. Fei-Fei, “Perceptual losses for real-time style transfer and super-resolution,” in Proc. ECCV . Springer, 2016, pp. 694–711
2016
Earlier work this paper cites.
J. S. Chung and A. Zisserman, “Out of time: automated lip sync in the wild,” in Proc. ACCV . Springer, 2016, pp. 251–263
2016
Earlier work this paper cites.
J. S. Chung, A. Jamaludin, and A. Zisserman, “You said that?” in Proc. BMVC , 2017
2017
Earlier work this paper cites.
S. Suwajanakorn, S. M. Seitz, and I. Kemelmacher-Shlizerman, “Synthesizing obama: learning lip sync from audio,” ACM TOG , vol. 36, no. 4, pp. 1–13, 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. NeurIPS , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
Y. M. Assael, B. Shillingford, S. Whiteson, and N. de Freitas, “Lipnet: End-to-end sentence-level lipreading,” GPU Technology Conference , 2017. [Online]. Available: https://github.com/Fengdalu/LipNet-PyTorch
2017
Earlier work this paper cites.
L. Chen, Z. Li, R. K. Maddox, Z. Duan, and C. Xu, “Lip movements generation at a glance,” in Proc. ECCV , 2018, pp. 520–535
2018
Earlier work this paper cites.
O. Wiles, A. Koepke, and A. Zisserman, “X2face: A network for controlling face generation using images, audio, and pose codes,” in Proc. ECCV , 2018, pp. 670–686
2018
Earlier work this paper cites.
H. Kim, P. Garrido, A. Tewari, W. Xu, J. Thies, M. Niessner, P. Pérez, C. Richardt, M. Zollhöfer, and C. Theobalt, “Deep video portraits,” ACM TOG , vol. 37, no. 4, pp. 1–14, 2018
2018
Earlier work this paper cites.
K. Nagano, J. Seo, J. Xing, L. Wei, Z. Li, S. Saito, A. Agarwal, J. Fursund, and H. Li, “pagan: real-time avatars using dynamic textures,” ACM TOG , vol. 37, no. 6, pp. 1–12, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
W.-S. Lai, J.-B. Huang, O. Wang, E. Shechtman, E. Yumer, and M.-H. Yang, “Learning blind video temporal consistency,” in Proc. ECCV , 2018, pp. 170–185
2018
Cited alongside, same era.
M. Wang, G.-Y. Yang, J.-K. Lin, S.-H. Zhang, A. Shamir, S.-P. Lu, and S.-M. Hu, “Deep online video stabilization with multi-grid warping transformation learning,” TIP , vol. 28, no. 5, pp. 2283–2292, 2018
2018
Cited alongside, same era.
L. Chen, R. K. Maddox, Z. Duan, and C. Xu, “Hierarchical cross-modal talking face generation with dynamic pixel-wise loss,” in Proc. CVPR , 2019, pp. 7832–7841
2019
Cited alongside, same era.
O. Fried, A. Tewari, M. Zollhöfer, A. Finkelstein, E. Shechtman, D. B. Goldman, K. Genova, Z. Jin, C. Theobalt, and M. Agrawala, “Text-based editing of talking-head video,” ACM TOG , vol. 38, no. 4, pp. 1–14, 2019
2019
Cited alongside, same era.
H. Zhou, Y. Sun, W. Wu, C. C. Loy, X. Wang, and Z. Liu, “Pose-controllable talking face generation by implicitly modularized audio-visual representation,” in Proc. CVPR , 2021, pp. 4176–4186
2021
Later among the works it cites.
H. Wu, J. Jia, H. Wang, Y. Dou, C. Duan, and Q. Deng, “Imitating arbitrary talking style for realistic audio-driven talking face synthesis,” in Proc. ACM MM , 2021, pp. 1478–1486
2021
Later among the works it cites.
X. Ji, H. Zhou, K. Wang, W. Wu, C. C. Loy, X. Cao, and F. Xu, “Audio-driven emotional video portraits,” in Proc. CVPR , 2021, pp. 14 080–14 089
2021
Later among the works it cites.
X. Yao, O. Fried, K. Fatahalian, and M. Agrawala, “Iterative text-based editing of talking-heads using neural retargeting,” ACM TOG , vol. 40, no. 3, pp. 1–14, 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Zhou, Y. Liu, Z. Liu, P. Luo, and X. Wang, “Talking face generation by adversarially disentangled audio-visual representation,” in Proc. AAAI , vol. 33, no. 01, 2019, pp. 9299–9306
2019
Cited alongside, same era.
A. Siarohin, S. Lathuilière, S. Tulyakov, E. Ricci, and N. Sebe, “First order motion model for image animation,” in Proc. NeurIPS , 2019, pp. 7137–7147
2019
Cited alongside, same era.
T.-C. Wang, M.-Y. Liu, A. Tao, G. Liu, B. Catanzaro, and J. Kautz, “Few-shot video-to-video synthesis,” in Proc. NeurIPS , 2019, pp. 5013–5024
2019
Cited alongside, same era.
E. Zakharov, A. Shysheya, E. Burkov, and V. Lempitsky, “Few-shot adversarial learning of realistic neural talking head models,” in Proc. ICCV , 2019, pp. 9459–9468
2019
Cited alongside, same era.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech: Fast, robust and controllable text to speech,” 2019, pp. 3171–3180
2019
Cited alongside, same era.
Y. Zhou, X. Han, E. Shechtman, J. Echevarria, E. Kalogerakis, and D. Li, “Makelttalk: speaker-aware talking-head animation,” ACM TOG , vol. 39, no. 6, pp. 1–15, 2020
2020
Cited alongside, same era.
K. Prajwal, R. Mukhopadhyay, V. P. Namboodiri, and C. Jawahar, “A lip sync expert is all you need for speech to lip generation in the wild,” in Proc. ACM MM , 2020, pp. 484–492
2020
Cited alongside, same era.
K. Vougioukas, S. Petridis, and M. Pantic, “Realistic speech-driven facial animation with gans,” IJCV , vol. 128, no. 5, pp. 1398–1413, 2020
2020
Cited alongside, same era.
2021
Later among the works it cites.
Y. Lu, J. Chai, and X. Cao, “Live Speech Portraits: Real-time photorealistic talking-head animation,” ACM Transactions on Graphics , vol. 40, no. 6, December 2021
2021
Later among the works it cites.
A. Lahiri, V. Kwatra, C. Frueh, J. Lewis, and C. Bregler, “Lipsync3d: Data-efficient learning of personalized 3d talking faces from video using pose and lighting normalization,” in Proc. CVPR , 2021, pp. 2755–2764
2021
Later among the works it cites.
L. Song, B. Liu, G. Yin, X. Dong, Y. Zhang, and J.-X. Bai, “Tacr-net: Editing on deep video and voice portraits,” in Proc. ACM MM , 2021, pp. 478–486
2021
Later among the works it cites.
C. Zhang, Y. Zhao, Y. Huang, M. Zeng, S. Ni, M. Budagavi, and X. Guo, “Facial: Synthesizing dynamic talking face with implicit attribute learning,” in Proc. ICCV , 2021, pp. 3867–3876
2021
Later among the works it cites.
Y. Feng, H. Feng, M. J. Black, and T. Bolkart, “Learning an animatable detailed 3d face model from in-the-wild images,” ACM TOG , vol. 40, no. 4, pp. 1–13, 2021
2021
Later among the works it cites.
M. C. Doukas, S. Zafeiriou, and V. Sharmanska, “Headgan: One-shot neural head synthesis and editing,” in Proc. ICCV , 2021, pp. 14 398–14 407
2021
Later among the works it cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR , 2021
2021
Later among the works it cites.
R. Liu, H. Deng, Y. Huang, X. Shi, L. Lu, W. Sun, X. Wang, J. Dai, and H. Li, “Fuseformer: Fusing fine-grained information in transformers for video inpainting,” in Proc. ICCV , 2021, pp. 14 040–14 049
2021
Later among the works it cites.
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable detr: Deformable transformers for end-to-end object detection,” in ICLR , 2021
2021
Later among the works it cites.
Y. Zeng, H. Yang, H. Chao, J. Wang, and J. Fu, “Improving visual quality of image synthesis by a token-based generator with transformers,” in Proc. NeurIPS , 2021, pp. 21 125–21 137
2021
Later among the works it cites.
Z. Guo, D. Guo, H. Zheng, Z. Gu, B. Zheng, and J. Dong, “Image harmonization with transformer,” in Proc. ICCV , 2021, pp. 14 870–14 879
2021
Later among the works it cites.
E. Wood, T. Baltrušaitis, C. Hewitt, S. Dziadzio, T. J. Cashman, and J. Shotton, “Fake it till you make it: Face analysis in the wild using synthetic data alone,” in Proc. ICCV , 2021, pp. 3681–3691
2021
Later among the works it cites.
A. Tang, Y. Huang, J. Ling, Z. Zhang, Y. Zhang, R. Xie, and L. Song, “Generative compression for face video: A hybrid scheme,” in Proc. ICME , 2022
2022
Closest in time.
Y. Fan, Z. Lin, J. Saito, W. Wang, and T. Komura, “Faceformer: Speech-driven 3d facial animation with transformers,” in Proc. CVPR , 2022, pp. 18 770–18 780
2022
Closest in time.
L. Chen, Z. Wu, J. Ling, R. Li, X. Tan, and S. Zhao, “Transformer-s2a: Robust and efficient speech-to-animation,” in Proc. ICASSP . IEEE, 2022, pp. 7247–7251
2022
Closest in time.
J. Ren, Q. Zheng, Y. Zhao, X. Xu, and C. Li, “Dlformer: Discrete latent transformer for video inpainting,” in Proc. CVPR , 2022, pp. 3511–3520
2022
Closest in time.
C. Li, L. Song, S. Chen, R. Xie, and W. Zhang, “Deep online video stabilization using imu sensors,” TMM , 2022
2022
Closest in time.