Fetching the paper…
Reading the bibliography…
In this paper, we present a generic and robust multimodal synthesis system that produces highly natural speech and facial expression simultaneously.
R. E. Crociere and L. R. Rabiner, “Multirate digital signal processing,” Prentice Hall, Englewood Cliffs
1983
Earlier work this paper cites.
T. Nguyen, “Near-perfect-reconstruction pseudo-qmf banks,” IEEE Transactions on Signal Processing, Vol. 42, No.1,
1994
Earlier work this paper cites.
A. J. Hunt and A. W. Black, “Unit selection in a concatenative speech synthesis system using a large speech database,” in 1996 IEEE International Conference on Acoustics, Speech, and Signal Processing Conference Proceedings
1996
Earlier work this paper cites.
A. W. Black and P. A. Taylor, “Automatically clustering similar units for unit selection in speech synthesis.,” 1997
1997
Earlier work this paper cites.
K. Tokuda, T. Yoshimura, T. Masuko, T. Kobayashi, and T. Kitamura, “Speech parameter generation algorithms for hmm-based speech synthesis,” in 2000 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No. 00CH37100)
2000
Earlier work this paper cites.
H. Zen, K. Tokuda, and A. W. Black, “Statistical parametric speech synthesis,” speech communication
2009
Earlier work this paper cites.
H. Zen, A. Senior, and M. Schuster, “Statistical parametric speech synthesis using deep neural networks,” in 2013 ieee international conference on acoustics, speech and signal processing
2013
Earlier work this paper cites.
Y. Fan, Y. Qian, F. Xie, and F. K. Soong, “TTS Synthesis with Bidirectional LSTM based Recurrent Neural Networks,” INTERSPEECH
2014
Earlier work this paper cites.
C. Chen, Y. Weng, S. Zhou, Y. Tong, and K. Zhou, “Facewarehouse: a 3D facial expression database for visual computing,” IEEE Transactions on Visualization and Computer Graphics
2014
Earlier work this paper cites.
O. Vinyals, Ł. Kaiser, T. Koo, S. Petrov, I. Sutskever, and G. Hinton, “Grammar as a foreign language,” in Advances in neural information processing systems
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Cited alongside, same era.
2017
Cited alongside, same era.
J. Sotelo, S. Mehri, K. Kumar, J. F. Santos, K. Kastner, A. Courville, and Y. Bengio, “Char2wav: End-to-end speech synthesis,” 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
W.-N. Hsu, Y. Zhang, R. J. Weiss, H. Zen, Y. Wu, Y. Wang, Y. Cao, Y. Jia, Z. Chen, J. Shen, et al
2018
Later among the works it cites.
T. Okamoto, T. Toda, Y. Shiga, and H. Kawai, “Improving FFTNet vocoder with noise shaping and subband approaches,” in 2018 IEEE Spoken Language Technology Workshop (SLT)
2018
Later among the works it cites.
T. Okamoto, K. Tachibana, T. Toda, Y. Shiga, and H. Kawai, “An investigation of subband WaveNet vocoder covering entire audible frequency range with limited acoustic features,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
2018
Later among the works it cites.
H. Kim, P. Garrido, A. Tewari, W. Xu, J. Thies, M. Nießner, P. Perez, C. Richardt, M. Zollhöfer, and C. Theobalt, “Deep video portraits,” Siggraph
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, et al
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
T. Okamoto, K. Tachibana, T. Toda, Y. Shiga, and H. Kawai, “An investigation of subband wavenet vocoder covering entire audible frequency range with limited acoustic features,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
2018
Later among the works it cites.
2019
Closest in time.
C. Yu, H. Lu, and D. Yu, “Duration Informed Attention Network For Text-To-Speech Analysis,” U.S. Provisional application, Pending
2019
Closest in time.
W.-N. Hsu, Y. Zhang, R. J. Weiss, Y.-A. Chung, Y. Wang, Y. Wu, and J. Glass, “Disentangling correlated speaker and noise for speech synthesis via data augmentation and adversarial factorization,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
2019
Closest in time.
J.-M. Valin and J. Skoglund, “Lpcnet: Improving neural speech synthesis through linear prediction,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
2019
Closest in time.
O. Fried, A. Tewari, M. Zollhöfer, A. Finkelstein, E. Shechtman, D. B. Goldman, K. Genova, Z. Jin, C. Theobalt, and M. Agrawala, “Text-based editing of talking-head video,” ACM Transactions on Graphics
2019
Closest in time.