Fetching the paper…
Reading the bibliography…
In this work, we extend ClariNet (Ping et al., 2019), a fully end-to-end speech synthesis model (i.e., text-to-wave), to generate high-fidelity speech from multiple speakers.
CrowdMOS: An approach for crowdsourcing mean opinion score studies
F. Ribeiro, D. Florêncio, C. Zhang, and M. Seltzer · 2011
Earlier work this paper cites.
Statistical parametric speech synthesis using deep neural networks
H. Zen, A. Senior, and M. Schuster · 2013
Earlier work this paper cites.
Multi-speaker modeling and speaker adaptation for DNN-based TTS synthesis
Y. Fan, Y. Qian, F. K. Soong, and L. He · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Deep learning for acoustic modeling in parametric speech generation: A systematic review of existing techniques and future trends
Z.-H. Ling, S.-Y. Kang, H. Zen, A. Senior, M. Schuster, X.-J. Qian, H. M. Meng, and L. Deng · 2015
Earlier work this paper cites.
A study of speaker adaptation for DNN-based speech synthesis
Z. Wu, P. Swietojanski, C. Veaux, S. Renals, and S. King · 2015
Earlier work this paper cites.
WaveNet: A generative model for raw audio
A. v. d. Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Deep neural network-based speaker embeddings for end-to-end speaker verification
D. Snyder, P. Ghahremani, D. Povey, D. Garcia-Romero, Y. Carmiel, and S. Khudanpur · 2016
Cited alongside, same era.
On the training of DNN-based average voice model for speech synthesis
S. Yang, Z. Wu, and L. Xie · 2016
Cited alongside, same era.
Deep Voice 2: Multi-speaker neural text-to-speech
S. Arik, G. Diamos, A. Gibiansky, J. Miller, K. Peng, W. Ping, J. Raiman, and Y. Zhou · 2017
Cited alongside, same era.
Deep Voice: Real-time neural text-to-speech
S. Ö. Arık, M. Chrzanowski, A. Coates, G. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, A. Ng, J. Raiman, S. Sengupta, and M. Shoeybi · 2017
Cited alongside, same era.
Char2Wav: End-to-end speech synthesis
J. Sotelo, S. Mehri, K. Kumar, J. F. Santos, K. Kastner, A. Courville, and Y. Bengio · 2017
Cited alongside, same era.
Tacotron: Towards end-to-end speech synthesis
Y. Wang, R. J. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. V. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous · 2017
Deep Voice 3: 2000-speaker neural text-to-speech
W. Ping, K. Peng, A. Gibiansky, S. O. Arik, A. Kannan, S. Narang, J. Raiman, and J. Miller · 2018
Later among the works it cites.
Natural TTS synthesis by conditioning WaveNet on Mel spectrogram predictions
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu · 2018
Later among the works it cites.
VoiceLoop: Voice fitting and synthesis via a phonological loop
Y. Taigman, L. Wolf, A. Polyak, and E. Nachmani · 2018
Later among the works it cites.
Parallel Neural Text-to-Speech
K. Peng, W. Ping, Z. Song, and K. Zhao · 2019
Closest in time.
ClariNet: Parallel Wave Generation in End-to-End Text-to-Speech
W. Ping, K. Peng, and J. Chen · 2019
Closest in time.
LibriTTS: A corpus derived from librispeech for text-to-speech
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Neural voice cloning with a few samples
S. O. Arik, J. Chen, K. Peng, W. Ping, and Y. Zhou · 2018
Cited alongside, same era.
Robust speaker-adaptive hmm-based text-to-speech synthesis
J. Yamagishi, T. Nose, H. Zen, Z. Ling, T. Toda, K. Tokuda, S. King, and S. Renals
Cited in the paper.
Robust speaker-adaptive HMM-based text-to-speech synthesis
J. Yamagishi, T. Nose, H. Zen, Z.-H. Ling, T. Toda, K. Tokuda, S. King, and S. Renals
Cited in the paper.
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu · 2019
Closest in time.