Y. Wang, R. J. Skerry-Ryan, D. Stanton et al. , “Tacotron: Towards end-to-end speech synthesis,” in Proc. ISCA Interspeech , 2017, pp. 4006–4010
2017
Cited alongside, same era.
X. Wang, S. Takaki, and J. Yamagishi, “An autoregressive recurrent mixture density network for parametric speech synthesis,” in Proc. IEEE ICASSP , 2017, pp. 4895–4899
2017
Cited alongside, same era.
K. Ito, “The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. NIPS , 2017, pp. 5998–6008
2017
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu, “Natural TTS synthesis by conditioning wavenet on MEL spectrogram predictions,” in Proc. IEEE ICASSP , 2018, pp. 4779–4783
2018
Cited alongside, same era.
N. Li, S. Liu, Y. Liu, S. Zhao, M. Liu, and M. Zhou, “Close to human quality TTS with transformer,” arXiv preprint arXiv:1809.08895 , 2018
Original
2018
Cited alongside, same era.
R. J. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. J. Weiss, R. Clark, and R. A. Saurous, “Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,” in Proc. ICML , 2018, pp. 4700–4709
2018
Cited alongside, same era.
Y. Wang, D. Stanton, Y. Zhang, R. J. Skerry-Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, and R. A. Saurous, “Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,” in Proc. ICML , 2018, pp. 5167–5176
2018
Cited alongside, same era.
K. Akuzawa, Y. Iwasawa, and Y. Matsuo, “Expressive speech synthesis via modeling expressions with variational autoencoder,” in Proc. ISCA Interspeech , 2018, pp. 3067–3071
2018
Cited alongside, same era.