Fetching the paper…
Reading the bibliography…
This paper proposes a hierarchical generative model with a multi-grained latent variable to synthesize expressive speech.
Y. Fan, Y. Qian, F.-L. Xie, and F. K. Soong, “TTS synthesis with bidirectional LSTM based recurrent neural networks,” in
1968
Earlier work this paper cites.
C. M. Bishop, “Mixture density networks,” Neural Computing Research Group, Aston University, Tech. Rep. NCRG/94/004, 1994
1994
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,”
1997
Earlier work this paper cites.
H. Zen, K. Tokuda, T. Masuko, T. Kobayasih, and T. Kitamura, “A hidden semi-Markov model-based speech synthesis system,”
2007
Earlier work this paper cites.
H. Zen, A. Senior, and M. Schuster, “Statistical parametric speech synthesis using deep neural networks,” in
2013
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes.” in
2014
Earlier work this paper cites.
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer, “Scheduled sampling for sequence prediction with recurrent neural networks,” in
2015
Earlier work this paper cites.
Z. Wu and S. King, “Investigating gated recurrent networks for speech synthesis,” in
2016
Earlier work this paper cites.
M. Morise, F. Yokomori, and K. Ozawa, “WORLD: a vocoder-based high-quality speech synthesis system for real-time applications,”
2016
Earlier work this paper cites.
K. Hashimoto, K. Oura, Y. Nankaku, and K. Tokuda, “Trajectory training considering global variance for speech synthesis based on neural networks,” in
2016
Cited alongside, same era.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio
2017
Cited alongside, same era.
2017
Cited alongside, same era.
K. Inoue, S. Hara, M. Abe, N. Hojo, and Y. Ijima, “An investigation to transplant emotional expressions in DNN-based TTS synthesis,” in
2017
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan
2018
Cited alongside, same era.
2018
Later among the works it cites.
N. Li, S. Liu, Y. Liu, S. Zhao, and M. Liu, “Neural speech synthesis with transformer network,” in
2019
Later among the works it cites.
Y.-J. Zhang, S. Pan, L. He, and Z.-H. Ling, “Learning latent representations for style control and transfer in end-to-end speech synthesis,” in
2019
Later among the works it cites.
Y. Lee and T. Kim, “Robust and fine-grained prosody control of end-to-end speech synthesis,” in
2019
Later among the works it cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
2018
Cited alongside, same era.
K. Akuzawa, Y. Iwasawa, and Y. Matsuo, “Expressive speech synthesis via modeling expressions with variational autoencoder,” in
2018
Cited alongside, same era.
“Speech signal processing toolkit (SPTK),”
Cited in the paper.
“Open JTalk,”
Cited in the paper.
T. Hayashi, S. Watanabe, T. Toda, K. Takeda, S. Toshniwal, and K. Livescu, “Pre-trained text embeddings for enhanced text-to-speech synthesis,” in
2019
Later among the works it cites.
G. Sun, Y. Zhang, R. J. Weiss, Y. Cao, H. Zen, and Y. Wu, “Fully-hierarchical fine-grained prosody modeling for interpretable speech synthesis,” in
2020
Closest in time.
G. Sun, Y. Zhang, R. J. Weiss, Y. Cao, H. Zen, A. Rosenberg, B. Ramabhadran, and Y. Wu, “Generating diverse and natural text-to-speech samples using a quantized fine-grained vae and autoregressive prosody prior,” in
2020
Closest in time.