Fetching the paper…
Reading the bibliography…
Talking head generation is to synthesize a lip-synchronized talking head video by inputting an arbitrary face image and corresponding audio clips.
S. E. Eskimez, R. K. Maddox, C. Xu, Z. Duan, End-to-end generation of talking faces from noisy speech, in: 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2020, pp. 1948–1952
1952
Earlier work this paper cites.
P. Ekman, W. Friesen, Facial action coding system: A technique for the measurement of facial movement, in: Consulting Psychologists, Consulting Psychologists Press, 1978
1978
Earlier work this paper cites.
F. Ringeval, E. Marchi, M. Mehu, K. Scherer, B. Schuller, Face reading from speech—predicting facial action units from audio cues, in: Sixteenth Annual Conference of the International Speech Communication Association (INTERSPEECH), 2015, pp. 1977–1981
1981
Earlier work this paper cites.
C. Bregler, M. Covell, M. Slaney, Video rewrite: Driving visual speech with audio, in: Proceedings of the 24th Annual Conference on Computer graphics and Interactive Techniques, 1997, pp. 353–360
1997
Earlier work this paper cites.
E. Yamamoto, S. Nakamura, K. Shikano, Lip movement synthesis from speech based on hidden markov models, Speech Communication 26 (1-2) (1998) 105–115
1998
Earlier work this paper cites.
K. H. Choi, J.-N. Hwang, Baum-welch hidden markov model inversion for reliable audio-to-visual conversion, in: 1999 IEEE Third Workshop on Multimedia Signal Processing, IEEE, 1999, pp. 175–180
1999
Earlier work this paper cites.
Z. Wang, A. C. Bovik, H. R. Sheikh, E. P. Simoncelli, Image quality assessment: from error visibility to structural similarity, IEEE Transactions on Image Processing 13 (4) (2004) 600–612
2004
Earlier work this paper cites.
M. Cooke, J. Barker, S. Cunningham, X. Shao, An audio-visual corpus for speech perception and automatic speech recognition, The Journal of the Acoustical Society of America 120 (5) (2006) 2421–2424
2006
Earlier work this paper cites.
L. Xie, Z. Liu, A coupled HMM approach to video-realistic speech animation, Pattern Recognition 40 (8) (2007) 2325–2340
2007
Earlier work this paper cites.
D. E. King, Dlib-ml: A machine learning toolkit, The Journal of Machine Learning Research 10 (2009) 1755–1758
2009
Earlier work this paper cites.
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, Y. Bengio, Generative adversarial networks, in: Advances in Neural Information Processing Systems (NIPS), 2014, pp. 2672–2680
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
B. Fan, L. Wang, F. K. Soong, L. Xie, Photo-real talking head with deep bidirectional lstm, in: 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2015, pp. 4884–4888
2015
Earlier work this paper cites.
O. Ronneberger, P. Fischer, T. Brox, U-net: Convolutional networks for biomedical image segmentation, in: International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2015, pp. 234–241
2015
Earlier work this paper cites.
N. Harte, E. Gillen, TCD-TIMIT: An audio-visual corpus of continuous speech, IEEE Transactions on Multimedia 17 (5) (2015) 603–615
2015
Earlier work this paper cites.
T. Baltrušaitis, M. Mahmoud, P. Robinson, Cross-dataset learning and person-specific normalisation for automatic action unit detection, in: 2015 11th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG), Vol. 6, IEEE, 2015, pp. 1–6
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al., Imagenet large scale visual recognition challenge, International Journal of Computer Vision (IJCV) 115 (3) (2015) 211–252
2015
Earlier work this paper cites.
J. S. Chung, A. Zisserman, Out of time: automated lip sync in the wild, in: Asian conference on computer vision (ACCV), 2016, pp. 251–263
2016
Earlier work this paper cites.
J. Yan, W. Zheng, Q. Xu, G. Lu, H. Li, B. Wang, Sparse kernel reduced-rank regression for bimodal emotion recognition from facial expression and speech, IEEE Transactions on Multimedia 18 (7) (2016) 1319–1329
2016
Earlier work this paper cites.
F. Milletari, N. Navab, S. A. Ahmadi, V-net: Fully convolutional neural networks for volumetric medical image segmentation, in: 2016 Fourth International Conference on 3D Vision (3DV), IEEE, 2016, pp. 565–571
2016
Earlier work this paper cites.
J. Johnson, A. Alahi, L. Fei-Fei, Perceptual losses for real-time style transfer and super-resolution, in: European Conference on Computer Vision (ECCV), 2016, pp. 694–711
2016
Earlier work this paper cites.
S. Suwajanakorn, S. M. Seitz, I. Kemelmacher-Shlizerman, Synthesizing obama: learning lip sync from audio, ACM Trans. Graph. 36 (4) (2017) 95:1–95:13
2017
Earlier work this paper cites.
J. S. Chung, A. Jamaludin, A. Zisserman, You said that?, in: British Machine Vision Conference (BMVC), 2017, pp. 1–12
2017
Earlier work this paper cites.
Z. Meng, S. Han, Y. Tong, Listen to your face: Inferring facial action units from audio channel, IEEE Transactions on Affective Computing 10 (4) (2017) 537–551
2017
Cited alongside, same era.
N. Takahashi, M. Gygli, L. Van Gool, Aenet: Learning deep audio features for video analysis, IEEE Transactions on Multimedia 20 (3) (2017) 513–524
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, I. Polosukhin, Attention is all you need, in: Advances in neural information processing systems, 2017, pp. 5998–6008
2017
Cited alongside, same era.
B. Martinez, M. F. Valstar, B. Jiang, M. Pantic, Automatic analysis of facial actions: A survey, IEEE Transactions on Affective Computing 10 (3) (2017) 325–347
2017
Cited alongside, same era.
A. Jamaludin, J. S. Chung, A. Zisserman, You said that?: Synthesising talking faces from audio, International Journal of Computer Vision 127 (11) (2019) 1767–1779
2019
Later among the works it cites.
N. Liu, T. Zhou, Y. Ji, Z. Zhao, L. Wan, Synthesizing talking faces from text and audio: an autoencoder and sequence-to-sequence convolutional neural network, Pattern Recognition 102 (2020) 107231
2020
Later among the works it cites.
2020
Later among the works it cites.
Y. Zhou, X. Han, E. Shechtman, J. Echevarria, E. Kalogerakis, D. Li, Makelttalk: speaker-aware talking-head animation, ACM Transactions on Graphics (TOG) 39 (6) (2020) 1–15
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Zhou, B. E. Shi, Photorealistic facial expression synthesis by the conditional difference adversarial autoencoder, in: 2017 Seventh International Conference on Affective Computing and Intelligent Interaction (ACII), IEEE, 2017, pp. 370–376
2017
Cited alongside, same era.
L. Chen, Z. Li, R. K. Maddox, Z. Duan, C. Xu, Lip movements generation at a glance, in: Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 520–535
2018
Cited alongside, same era.
2018
Cited alongside, same era.
A. Pumarola, A. Agudo, A. M. Martinez, A. Sanfeliu, F. Moreno-Noguer, Ganimation: Anatomically-aware facial animation from a single image, in: Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 818–833
2018
Cited alongside, same era.
T. Baltrušaitis, A. Zadeh, Y. C. Lim, L.-P. Morency, Openface 2.0: Facial behavior analysis toolkit, in: 2018 13th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2018), IEEE, 2018, pp. 59–66
2018
Cited alongside, same era.
Z. Shao, Z. Liu, J. Cai, L. Ma, Deep adaptive attention for joint facial action unit detection and face alignment, in: Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 705–720
2018
Cited alongside, same era.
N. E. D. Elmadany, Y. He, L. Guan, Multimodal learning for human action recognition via bimodal/multimodal hybrid centroid canonical correlation analysis, IEEE Transactions on Multimedia 21 (5) (2018) 1317–1331
2018
Cited alongside, same era.
Y. Zhang, K. Li, K. Li, L. Wang, B. Zhong, Y. Fu, Image super-resolution using very deep residual channel attention networks, in: Proceedings of the European conference on computer vision (ECCV), 2018, pp. 286–301
2018
Cited alongside, same era.
L. Yu, J. Yu, M. Li, Q. Ling, Multimodal inputs driven talking face generation with spatial–temporal dependency, IEEE Transactions on Circuits and Systems for Video Technology 31 (1) (2020) 203–216
2020
Later among the works it cites.
Z. Liu, D. Liu, Y. Wu, Region based adversarial synthesis of facial action units, in: International Conference on Multimedia Modeling, 2020, pp. 514–526
2020
Later among the works it cites.
K. Prajwal, R. Mukhopadhyay, V. P. Namboodiri, C. Jawahar, A lip sync expert is all you need for speech to lip generation in the wild, in: Proceedings of the 28th ACM International Conference on Multimedia, 2020, pp. 484–492
2020
Later among the works it cites.
L. Chen, G. Cui, C. Liu, Z. Li, Z. Kou, Y. Xu, C. Xu, Talking-head generation with rhythmic head motion, in: European Conference on Computer Vision, Springer, 2020, pp. 35–51
2020
Later among the works it cites.
L. Pang, S. Zhu, C.-W. Ngo, Deep multimodal learning for affective analysis and retrieval, IEEE Transactions on Multimedia 17 (11) (2015) 2008–2020
2020
Later among the works it cites.
C. Zhang, Z. Yang, X. He, L. Deng, Multimodal intelligence: Representation learning, information fusion, and applications, IEEE Journal of Selected Topics in Signal Processing 14 (3) (2020) 478–493
2020
Later among the works it cites.
P. Li, X. Li, Multimodal fusion with co-attention mechanism, in: 2020 IEEE 23rd International Conference on Information Fusion (FUSION), IEEE, 2020, pp. 1–8
2020
Later among the works it cites.
Q. Wang, B. Wu, P. Zhu, P. Li, W. Zuo, Q. Hu, Eca-net: Efficient channel attention for deep convolutional neural networks, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 11531–11539
2020
Later among the works it cites.
J. Liu, Z. Liu, L. Wang, Y. Gao, L. Guo, J. Dang, Temporal attention convolutional network for speech emotion recognition with latent representation., in: INTERSPEECH, 2020, pp. 2337–2341
2020
Later among the works it cites.
R. Wu, G. Zhang, S. Lu, T. Chen, Cascade ef-gan: Progressive facial expression editing with local focuses, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 5021–5030
2020
Later among the works it cites.
Y. Guo, X. Zhang, X. Wu, Deep multi-modality soft-decoding of very low bit-rate face videos, in: Proceedings of the 28th ACM International Conference on Multimedia, 2020, pp. 3947–3955
2020
Later among the works it cites.
J. Tang, Z. Shao, L. Ma, Fine-grained expression manipulation via structured latent space, in: 2020 IEEE International Conference on Multimedia and Expo (ICME), IEEE, 2020, pp. 1–6
2020
Later among the works it cites.
Y. Sun, H. Zhou, Z. Liu, H. Koike, Speech2talking-face: Inferring and driving a face with synchronized audio-visual representation, in: Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence (IJCAI-21), 2021, pp. 1018–1024
2021
Later among the works it cites.
2021
Later among the works it cites.
H. Zhou, Y. Sun, W. Wu, C. C. Loy, X. Wang, Z. Liu, Pose-controllable talking face generation by implicitly modularized audio-visual representation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 4176–4186
2021
Later among the works it cites.
S. E. Eskimez, Y. Zhang, Z. Duan, Speech driven talking face generation from a single image and an emotion condition, IEEE Transactions on Multimedia (2021)
2021
Later among the works it cites.
X. Ji, H. Zhou, K. Wang, W. Wu, C. C. Loy, X. Cao, F. Xu, Audio-driven emotional video portraits, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 14080–14089
2021
Later among the works it cites.
J. Liu, S. Chen, L. Wang, Z. Liu, Y. Fu, L. Guo, J. Dang, Multimodal emotion recognition with capsule graph convolutional based representation fusion, in: ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021, pp. 6339–6343
2021
Later among the works it cites.