Fetching the paper…
Reading the bibliography…
Automatically generating videos in which synthesized speech is synchronized with lip movements in a talking head has great potential in many human-computer interaction scenarios.
D. Griffin and J. Lim, “Signal estimation from modified short-time fourier transform,” IEEE Transactions on acoustics, speech, and signal processing , vol. 32, no. 2, pp. 236–243, 1984
1984
Earlier work this paper cites.
O. Schreer, R. Englert, P. Eisert, and R. Tanger, “Real-time vision and speech driven avatars for multimedia applications,” IEEE Transactions on Multimedia , vol. 10, no. 3, pp. 352–360, 2008
2008
Earlier work this paper cites.
J. Yamagishi, T. Nose, H. Zen, Z.-H. Ling, T. Toda, K. Tokuda, S. King, and S. Renals, “Robust speaker-adaptive hmm-based text-to-speech synthesis,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 17, no. 6, pp. 1208–1230, 2009
2009
Earlier work this paper cites.
D. E. King, “Dlib-ml: A machine learning toolkit,” The Journal of Machine Learning Research , vol. 10, pp. 1755–1758, 2009
2009
Earlier work this paper cites.
A. Segal, D. Haehnel, and S. Thrun, “Generalized-icp.” in Robotics: science and systems , vol. 2, no. 4. Seattle, WA, 2009, p. 435
2009
Earlier work this paper cites.
A. L. Maas, A. Y. Hannun, and A. Y. Ng, “Rectifier nonlinearities improve neural network acoustic models,” in Proc. icml , vol. 30, no. 1. Citeseer, 2013, p. 3
2013
Earlier work this paper cites.
P. Garrido, L. Valgaerts, H. Sarmadi, I. Steiner, K. Varanasi, P. Perez, and C. Theobalt, “Vdub: Modifying face video of actors for plausible visual alignment to a dubbed audio track,” in Computer graphics forum , vol. 34, no. 2. Wiley Online Library, 2015, pp. 193–204
2015
Earlier work this paper cites.
Y. Fan, Y. Qian, F. K. Soong, and L. He, “Multi-speaker modeling and speaker adaptation for dnn-based tts synthesis,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2015, pp. 4475–4479
2015
Earlier work this paper cites.
F. Schroff, D. Kalenichenko, and J. Philbin, “Facenet: A unified embedding for face recognition and clustering,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 815–823
2015
Earlier work this paper cites.
J. Charles, D. Magee, and D. Hogg, “Virtual immortality: Reanimating characters from tv shows,” in European Conference on Computer Vision . Springer, 2016, pp. 879–886
2016
Earlier work this paper cites.
B. Fan, L. Xie, S. Yang, L. Wang, and F. K. Soong, “A deep bidirectional lstm approach for video-realistic talking head,” Multimedia Tools and Applications , vol. 75, no. 9, pp. 5287–5309, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
S. Yang, Z. Wu, and L. Xie, “On the training of dnn-based average voice model for speech synthesis,” in 2016 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA) . IEEE, 2016, pp. 1–6
2016
Earlier work this paper cites.
P. Edwards, C. Landreth, E. Fiume, and K. Singh, “Jali: an animator-centric viseme model for expressive lip synchronization,” ACM Transactions on Graphics (TOG) , vol. 35, no. 4, pp. 1–11, 2016
2016
Earlier work this paper cites.
J. S. Chung and A. Zisserman, “Lip reading in the wild,” in Asian Conference on Computer Vision . Springer, 2016, pp. 87–103
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. Sotelo, S. Mehri, K. Kumar, J. F. Santos, K. Kastner, A. Courville, and Y. Bengio, “Char2wav: End-to-end speech synthesis,” 2017
2017
Earlier work this paper cites.
K. Ito and L. Johnson, “The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
2017
Earlier work this paper cites.
A. Bulat and G. Tzimiropoulos, “How far are we from solving the 2d & 3d face alignment problem?(and a dataset of 230,000 3d facial landmarks),” in Proceedings of the IEEE International Conference on Computer Vision , 2017, pp. 1021–1030
2017
Earlier work this paper cites.
S. Suwajanakorn, S. M. Seitz, and I. Kemelmacher-Shlizerman, “Synthesizing obama: learning lip sync from audio,” ACM Transactions on Graphics (TOG) , vol. 36, no. 4, pp. 1–13, 2017
2017
Earlier work this paper cites.
J. S. Chung, A. Jamaludin, and A. Zisserman, “You said that?” arXiv preprint arXiv:1705.02966 , 2017
2017
Earlier work this paper cites.
J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman, “Lip reading sentences in the wild,” in 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, 2017, pp. 3444–3453
2017
Cited alongside, same era.
2017
Cited alongside, same era.
P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros, “Image-to-image translation with conditional adversarial networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 1125–1134
2017
Cited alongside, same era.
N. Sadoughi Nourabadi, “Synthesizing naturalistic and meaningful speech-driven behaviors,” Ph.D. dissertation, 2017
2017
Cited alongside, same era.
A. Siarohin, S. Lathuilière, S. Tulyakov, E. Ricci, and N. Sebe, “First order motion model for image animation,” Advances in Neural Information Processing Systems , vol. 32, pp. 7137–7147, 2019
2019
Later among the works it cites.
W. Chae and Y. Kim, “Text-driven speech animation with emotion control,” KSII Transactions on Internet and Information Systems (TIIS) , vol. 14, no. 8, pp. 3473–3487, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
Y. Zhou, X. Han, E. Shechtman, J. Echevarria, E. Kalogerakis, and D. Li, “Makelttalk: speaker-aware talking-head animation,” ACM Transactions on Graphics (TOG) , vol. 39, no. 6, pp. 1–15, 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan et al. , “Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 4779–4783
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
E. Nachmani, A. Polyak, Y. Taigman, and L. Wolf, “Fitting new speakers based on a short untranscribed sample,” in International Conference on Machine Learning . PMLR, 2018, pp. 3683–3691
2018
Cited alongside, same era.
L. Yu, J. Yu, and Q. Ling, “Bltrcnn-based 3-d articulatory movement prediction: Learning articulatory synchronicity from both text and audio inputs,” IEEE Transactions on Multimedia , vol. 21, no. 7, pp. 1621–1632, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
P. Esser, E. Sutter, and B. Ommer, “A variational u-net for conditional appearance and shape generation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 8857–8866
2018
Cited alongside, same era.
C. Yu, H. Lu, N. Hu, M. Yu, C. Weng, K. Xu, P. Liu, D. Tuo, S. Kang, G. Lei et al. , “Durian: Duration informed attention network for speech synthesis,” Proc. Interspeech 2020 , pp. 2027–2031, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
E. Cooper, C.-I. Lai, Y. Yasuda, F. Fang, X. Wang, N. Chen, and J. Yamagishi, “Zero-shot multi-speaker text-to-speech with state-of-the-art neural speaker embeddings,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6184–6188
2020
Later among the works it cites.
X. Wen, M. Wang, C. Richardt, Z.-Y. Chen, and S.-M. Hu, “Photorealistic audio-driven video portraits,” IEEE Transactions on Visualization and Computer Graphics , vol. 26, no. 12, pp. 3457–3466, 2020
2020
Later among the works it cites.
R. Yi, Z. Ye, J. Zhang, H. Bao, and Y.-J. Liu, “Audio-driven talking face video generation with learning-based personalized head pose,” arXiv e-prints , pp. arXiv–2002, 2020
2020
Later among the works it cites.
J. Thies, M. Elgharib, A. Tewari, C. Theobalt, and M. Nießner, “Neural voice puppetry: Audio-driven facial reenactment,” in European Conference on Computer Vision . Springer, 2020, pp. 716–731
2020
Later among the works it cites.
L. Yu, J. Yu, M. Li, and Q. Ling, “Multimodal inputs driven talking face generation with spatial-temporal dependency,” IEEE Transactions on Circuits and Systems for Video Technology , 2020
2020
Later among the works it cites.
K. Prajwal, R. Mukhopadhyay, V. P. Namboodiri, and C. Jawahar, “A lip sync expert is all you need for speech to lip generation in the wild,” in Proceedings of the 28th ACM International Conference on Multimedia , 2020, pp. 484–492
2020
Later among the works it cites.
A. Nagrani, J. S. Chung, W. Xie, and A. Zisserman, “Voxceleb: Large-scale speaker verification in the wild,” Computer Speech & Language , vol. 60, p. 101027, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
E. Battenberg, R. Skerry-Ryan, S. Mariooryad, D. Stanton, D. Kao, M. Shannon, and T. Bagby, “Location-relative attention mechanisms for robust long-form speech synthesis,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6194–6198
2020
Later among the works it cites.
2020
Later among the works it cites.
2021
Closest in time.
G. Yang, S. Yang, K. Liu, P. Fang, W. Chen, and L. Xie, “Multi-band melgan: Faster waveform generation for high-quality text-to-speech,” in 2021 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2021, pp. 492–498
2021
Closest in time.
R. J. Weiss, R. Skerry-Ryan, E. Battenberg, S. Mariooryad, and D. P. Kingma, “Wave-tacotron: Spectrogram-free end-to-end text-to-speech synthesis,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 5679–5683
2021
Closest in time.
2021
Closest in time.
A. Richard, C. Lea, S. Ma, J. Gall, F. de la Torre, and Y. Sheikh, “Audio-and gaze-driven facial animation of codec avatars,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2021, pp. 41–50
2021
Closest in time.