“BS. 1534-1. method for the subjective assessment of intermediate sound quality (MUSHRA),”
I. Recommendation, · 2001
Earlier work this paper cites.
“Auto-encoding variational bayes,”
D. P. Kingma and M. Welling, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
D. P. Kingma and J. Ba, · 2015
Earlier work this paper cites.
“Wavenet: A generative model for raw audio,”
A. van den Oord, S. Dieleman, H. Zen, et al., · 2016
Earlier work this paper cites.
“Deep voice: Real-time neural text-to-speech,”
S. Ö. Arik, M. Chrzanowski, A. Coates, et al., · 2017
Earlier work this paper cites.
“Tacotron: A fully end-to-end text-to-speech synthesis model,”
Original
Y. Wang, R. J. Skerry-Ryan, D. Stanton, et al., · 2017
Earlier work this paper cites.
“Deep voice 2: Multi-speaker neural text-to-speech,”
A. Gibiansky, S. Ö. Arik, G. F. Diamos, et al., · 2017
Earlier work this paper cites.
“An investigation of multi-speaker training for wavenet vocoder,”
T. Hayashi, A. Tamamori, K. Kobayashi, K. Takeda, and T. Toda, · 2017
Earlier work this paper cites.
“Natural TTS synthesis by conditioning wavenet on MEL spectrogram predictions,”
J. Shen, R. Pang, R. J. Weiss, et al., · 2018
Earlier work this paper cites.