Fetching the paper…
Reading the bibliography…
Audio-driven talking face generation, which aims to synthesize talking faces with realistic facial animations (including accurate lip movements, vivid facial expression details and natural head poses) corresponding to the audio, has achieved rapid progress in recent years.
1997
Earlier work this paper cites.
J. Cassell, D. McNeill, K.-E. McCullough, Speech-gesture mismatches: Evidence for one underlying representation of linguistic and nonlinguistic information, Pragmatics & cognition 7 (1) (1999) 1–34
1999
Earlier work this paper cites.
P. Ekman, W. Friesen, J. Hager, Facial action coding system (facs) a human face, Salt Lake City (2002)
2002
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, Y. Bengio, Generative adversarial nets, in: Advances in neural information processing systems, 2014, pp. 2672–2680
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
B. Fan, L. Wang, F. K. Soong, L. Xie, Photo-real talking head with deep bidirectional lstm, in: 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2015, pp. 4884–4888
2015
Earlier work this paper cites.
J. S. Chung, A. Zisserman, Lip reading in the wild, in: Asian conference on computer vision, Springer, 2016, pp. 87–103
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, I. Polosukhin, Attention is all you need, Advances in neural information processing systems 30 (2017)
2017
Earlier work this paper cites.
S. Suwajanakorn, S. M. Seitz, I. Kemelmacher-Shlizerman, Synthesizing obama: learning lip sync from audio, ACM Transactions on Graphics (ToG) 36 (4) (2017) 1–13
2017
Earlier work this paper cites.
X. Mao, Q. Li, H. Xie, R. Y. Lau, Z. Wang, S. Paul Smolley, Least squares generative adversarial networks, in: Proceedings of the IEEE international conference on computer vision, 2017, pp. 2794–2802
2017
Earlier work this paper cites.
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, A. Lerer, Automatic differentiation in pytorch (2017)
2017
Earlier work this paper cites.
L. Chen, Z. Li, R. K. Maddox, Z. Duan, C. Xu, Lip movements generation at a glance, in: Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 520–535
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
O. Wiles, A. Sophia Koepke, A. Zisserman, X2face: A network for controlling face generation using images, audio, and pose codes, in: Proceedings of the European conference on computer vision (ECCV), 2018, pp. 670–686
2018
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, Improving language understanding by generative pre-training (2018)
2018
Earlier work this paper cites.
2018
Cited alongside, same era.
T. Baltrusaitis, A. Zadeh, Y. C. Lim, L.-P. Morency, Openface 2.0: Facial behavior analysis toolkit, in: 2018 13th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2018), IEEE, 2018, pp. 59–66
2018
Cited alongside, same era.
S. Tulyakov, M.-Y. Liu, X. Yang, J. Kautz, Mocogan: Decomposing motion and content for video generation, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 1526–1535
2018
Cited alongside, same era.
L. Chen, R. K. Maddox, Z. Duan, C. Xu, Hierarchical cross-modal talking face generation with dynamic pixel-wise loss, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 7832–7841
2019
Cited alongside, same era.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, S. Zagoruyko, End-to-end object detection with transformers, in: European conference on computer vision, Springer, 2020, pp. 213–229
2020
Later among the works it cites.
X. Wen, M. Wang, C. Richardt, Z.-Y. Chen, S.-M. Hu, Photorealistic audio-driven video portraits, IEEE Transactions on Visualization and Computer Graphics 26 (12) (2020) 3457–3466
2020
Later among the works it cites.
H. Zhou, Y. Sun, W. Wu, C. C. Loy, X. Wang, Z. Liu, Pose-controllable talking face generation by implicitly modularized audio-visual representation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 4176–4186
2021
Later among the works it cites.
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Zhou, Y. Liu, Z. Liu, P. Luo, X. Wang, Talking face generation by adversarially disentangled audio-visual representation, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33, 2019, pp. 9299–9306
2019
Cited alongside, same era.
E. Zakharov, A. Shysheya, E. Burkov, V. Lempitsky, Few-shot adversarial learning of realistic neural talking head models, in: Proceedings of the IEEE International Conference on Computer Vision, 2019, pp. 9459–9468
2019
Cited alongside, same era.
R. Girdhar, J. Carreira, C. Doersch, A. Zisserman, Video action transformer network, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 244–253
2019
Cited alongside, same era.
K. Prajwal, R. Mukhopadhyay, V. P. Namboodiri, C. Jawahar, A lip sync expert is all you need for speech to lip generation in the wild, in: Proceedings of the 28th ACM International Conference on Multimedia, 2020, pp. 484–492
2020
Cited alongside, same era.
L. Chen, G. Cui, C. Liu, Z. Li, Z. Kou, Y. Xu, C. Xu, Talking-head generation with rhythmic head motion, in: European Conference on Computer Vision, Springer, 2020, pp. 35–51
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Y. Zhou, X. Han, E. Shechtman, J. Echevarria, E. Kalogerakis, D. Li, Makelttalk: speaker-aware talking-head animation, ACM Transactions on Graphics (TOG) 39 (6) (2020) 1–15
2020
Cited alongside, same era.
K. Vougioukas, S. Petridis, M. Pantic, Realistic speech-driven facial animation with gans, International Journal of Computer Vision 128 (5) (2020) 1398–1413
2020
Cited alongside, same era.
C. Zhang, Y. Zhao, Y. Huang, M. Zeng, S. Ni, M. Budagavi, X. Guo, Facial: Synthesizing dynamic talking face with implicit attribute learning, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 3867–3876
2021
Later among the works it cites.
X. Ji, H. Zhou, K. Wang, W. Wu, C. C. Loy, X. Cao, F. Xu, Audio-driven emotional video portraits, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 14080–14089
2021
Later among the works it cites.
Y. Ren, G. Li, Y. Chen, T. H. Li, S. Liu, Pirenderer: Controllable portrait image generation via semantic neural rendering, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 13759–13768
2021
Later among the works it cites.
R. Li, S. Yang, D. A. Ross, A. Kanazawa, Learn to dance with aist++: Music conditioned 3d dance generation, arXiv e-prints (2021) arXiv–2101
2021
Later among the works it cites.
P. Esser, R. Rombach, B. Ommer, Taming transformers for high-resolution image synthesis, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 12873–12883
2021
Later among the works it cites.
E. Aksan, M. Kaufmann, P. Cao, O. Hilliges, A spatio-temporal transformer for 3d human motion prediction, in: 2021 International Conference on 3D Vision (3DV), IEEE, 2021, pp. 565–574
2021
Later among the works it cites.
2021
Later among the works it cites.
T.-C. Wang, A. Mallya, M.-Y. Liu, One-shot free-view neural talking-head synthesis for video conferencing, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 10039–10049
2021
Later among the works it cites.
B. Liang, Y. Pan, Z. Guo, H. Zhou, Z. Hong, X. Han, J. Han, J. Liu, E. Ding, J. Wang, Expressive talking head generation with granular audio-visual control, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 3387–3396
2022
Later among the works it cites.
L. Siyao, W. Yu, T. Gu, C. Lin, Q. Wang, C. Qian, C. C. Loy, Z. Liu, Bailando: 3d dance generation by actor-critic gpt with choreographic memory, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 11050–11059
2022
Later among the works it cites.