Fetching the paper…
Reading the bibliography…
Neural TTS has shown it can generate high quality synthesized speech.
“Average-voice based speech synthesis using HSMM-based speaker adaptation and adaptive training,”
J. Yamagishi and T. Kobayashi, · 2007
Earlier work this paper cites.
“Multispeaker modeling and speaker adaptation for DNN-based tts synthesis,”
Y. Fan, Y. Qian, F. K. Soong, and L. He, · 2015
Earlier work this paper cites.
“WaveNet: A generative model for raw audio,”
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, · 2016
Earlier work this paper cites.
“Multi-language multi-speaker acoustic modeling for LSTM-RNN based statistical parametric speech synthesis,”
B. Li and H. Zen, · 2016
Earlier work this paper cites.
“Tacotron: Towards end-to-end speech synthesis,”
Y. Wang, RJ Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, · 2017
Earlier work this paper cites.
“Char2Wav: End-to-end speech synthesis,”
J. Sotelo, S. Mehri, K. Kumar, J. F. Santos, K. Kastner, A. Courville, and Y. Bengio, · 2017
Earlier work this paper cites.
“Deep Voice: Real-time neural text-to-speech,”
S. O. Arik, M. Chrzanowski, A. Coates, G. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, A. Ng, J. Raiman, S. Sengupta, and M. Shoeybi, · 2017
Cited alongside, same era.
“Deep Voice 2: Multi-speaker neural text-to-speech,”
S. Arik, G. Diamos, A. Gibiansky, J. Miller, K. Peng, W. Ping, J. Raiman, and Y. Zhou, · 2017
Cited alongside, same era.
“Deep speaker: an end-to-end neural speaker embedding system,”
C. Li, X. Ma, B. Jiang, X. Li, X. Zhang, X. Liu, Y. Cao, A. Kannan, and Z. Zhu, · 2017
Cited alongside, same era.
“Natural TTS synthesis by conditioning WaveNet on Mel spectrogram predictions,”
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, RJ Skerry-Ryan, R. A. Saurous, Y.s Agiomyrgiannakis, and Y. Wu, · 2018
Cited alongside, same era.
“VoiceLoop: Voice fitting and synthesis via a phonological loop,”
Y. Taigman, L. Wolf, A. Polyak, and E. Nachmani, · 2018
Cited alongside, same era.
“Deep Voice 3: 2000-speaker neural text-to-speech,”
W. Ping, K. Peng, A. Gibiansky, S. O. Arik, A. Kannan, S. Narang, J. Raiman, and J. Miller, · 2018
Closest in time.
“Neural voice cloning with a few samples,”
S. O Arik, J. Chen, K. Peng, W. Ping, and Y. Zhou, · 2018
Closest in time.
“Fitting new speakers based on a short untranscribed sample,”
E. Nachmani, A. Polyak, Y. Taigman, and L. Wolf, · 2018
Closest in time.
“Transfer learning from speaker verification to multispeaker text-to-speech synthesis,”
Y. Jia, Y. Zhang, R. J. Weiss, Q. Wang, J. Shen, F. Ren, Z. Chen, P. Nguyen, R. Pang, I. L. Moreno, and Y. Wu, · 2018
Closest in time.
“Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,”
Y. Wang, D. Stanton, Y. Zhang, RJ Skerry-Ryan, E. Battenberg, J. Shor, Y. Xiao, F. Ren, Y. Jia, and R. A. Saurous, · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…