Fetching the paper…
Reading the bibliography…
Exploiting rich linguistic information in raw text is crucial for expressive text-to-speech (TTS).
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
2008
Earlier work this paper cites.
P. Taylor, Text-to-speech synthesis . Cambridge university press, 2009
2009
Earlier work this paper cites.
H. Zen, K. Tokuda, and A. W. Black, “Statistical parametric speech synthesis,” speech communication , vol. 51, no. 11, pp. 1039–1064, 2009
2009
Earlier work this paper cites.
S. Kübler, R. McDonald, and J. Nivre, “Dependency parsing,” Synthesis lectures on human language technologies , vol. 1, no. 1, pp. 1–127, 2009
2009
Earlier work this paper cites.
S. Takamichi, K. Kobayashi, K. Tanaka, T. Toda, and S. Nakamura, “The naist text-to-speech system for the blizzard challenge 2015,” in Proc. Blizzard Challenge workshop , vol. 2. Berlin, Germany, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2017
Earlier work this paper cites.
K. Ito, “The ljspeech dataset,” 2017. [Online]. Available: https://keithito.com/LJ-Speech-Dataset/
2017
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan et al. , “Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 4779–4783
2018
Earlier work this paper cites.
Y. Wang, D. Stanton, Y. Zhang, R.-S. Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, and R. A. Saurous, “Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,” in International Conference on Machine Learning . PMLR, 2018, pp. 5180–5189
2018
Cited alongside, same era.
2018
Cited alongside, same era.
M. Schlichtkrull, T. N. Kipf, P. Bloem, R. Van Den Berg, I. Titov, and M. Welling, “Modeling relational data with graph convolutional networks,” in European semantic web conference . Springer, 2018, pp. 593–607
2018
Cited alongside, same era.
Y.-J. Zhang, S. Pan, L. He, and Z.-H. Ling, “Learning latent representations for style control and transfer in end-to-end speech synthesis,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6945–6949
2020
Later among the works it cites.
T. Kenter, M. Sharma, and R. Clark, “Improving the prosody of rnn-based english text-to-speech synthesis by incorporating a bert model,” Proc. Interspeech 2020 , pp. 4412–4416, 2020
2020
Later among the works it cites.
Y. Xiao, L. He, H. Ming, and F. K. Soong, “Improving prosody with linguistic and bert derived features in multi-speaker based mandarin chinese neural tts,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6704–6708
2020
Later among the works it cites.
M. Zhang, “A survey of syntactic-semantic parsing based on constituent and dependency structures,” Science China Technological Sciences , pp. 1–23, 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
H. Guo, F. K. Soong, L. He, and L. Xie, “Exploiting syntactic features in a parsed tree to improve end-to-end tts,” Proc. Interspeech 2019 , pp. 4460–4464, 2019
2019
Cited alongside, same era.
T. Hayashi, S. Watanabe, T. Toda, K. Takeda, S. Toshniwal, and K. Livescu, “Pre-trained text embeddings for enhanced text-to-speech synthesis,” Proc. Interspeech 2019 , pp. 4430–4434, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
M. Wang, L. Yu, D. Zheng, Q. Gan, Y. Gai, Z. Ye, M. Li, J. Zhou, Q. Huang, C. Ma et al. , “Deep graph library: Towards efficient and scalable deep learning on graphs.” 2019
2019
Cited alongside, same era.
Databaker technology Inc. (Beijing), “Open source Chinese female voice database,” 2019. [Online]. Available: https://www.data-baker.com/open_source
2019
Cited alongside, same era.
A. Sun, J. Wang, N. Cheng, H. Peng, Z. Zeng, and J. Xiao, “Graphtts: graph-to-sequence modelling in neural text-to-speech,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6719–6723
2020
Cited alongside, same era.
“Bert-base-uncased model.” [Online]. Available: https://huggingface.co/bert-base-uncased/
Cited in the paper.
“Bert-base-chinese model.” [Online]. Available: https://huggingface.co/bert-base-chinese/
Cited in the paper.
2020
Later among the works it cites.
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” Advances in Neural Information Processing Systems , vol. 33, 2020
2020
Later among the works it cites.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech 2: Fast and high-quality end-to-end text to speech,” in International Conference on Learning Representations , 2021
2021
Closest in time.
2021
Closest in time.
C. Song, J. Li, Y. Zhou, Z. Wu, and H. Meng, “Syntactic representation learning for neural network based tts with syntactic parse tree traversal,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 6064–6068
2021
Closest in time.
Y.-J. Zhang and Z.-H. Ling, “Extracting and predicting word-level style variations for speech synthesis,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 1582–1593, 2021
2021
Closest in time.