Fetching the paper…
Reading the bibliography…
We introduce a technique for augmenting neural text-to-speech (TTS) with lowdimensional trainable speaker embeddings to generate different voices from a single model.
Speaker verification using adapted gaussian mixture models
D. A. Reynolds, T. F. Quatieri, and R. B. Dunn · 2000
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber · 2006
Earlier work this paper cites.
Robust speaker-adaptive hmm-based text-to-speech synthesis
J. Yamagishi, T. Nose, H. Zen, Z.-H. Ling, T. Toda, K. Tokuda, S. King, and S. Renals · 2009
Earlier work this paper cites.
Crowdmos: An approach for crowdsourcing mean opinion score studies
F. Ribeiro, D. Florêncio, C. Zhang, and M. Seltzer · 2011
Earlier work this paper cites.
Fast speaker adaptation of hybrid NN/HMM model for speech recognition based on discriminative learning of speaker code
O. Abdel-Hamid and H. Jiang · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Earlier work this paper cites.
Multi-speaker modeling and speaker adaptation for DNN-based TTS synthesis
Y. Fan, Y. Qian, F. K. Soong, and L. He · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
A study of speaker adaptation for DNN-based speech synthesis
Z. Wu, P. Swietojanski, C. Veaux, S. Renals, and S. King · 2015
Cited alongside, same era.
Unidirectional long short-term memory recurrent neural network with recurrent output layer for low-latency speech synthesis
H. Zen and H. Sak · 2015
Cited alongside, same era.
Neural architectures for named entity recognition
G. Lample, M. Ballesteros, K. Kawakami, S. Subramanian, and C. Dyer · 2016
Cited alongside, same era.
SampleRNN: An unconditional end-to-end neural audio generation model
S. Mehri, K. Kumar, I. Gulrajani, R. Kumar, S. Jain, J. Sotelo, A. Courville, and Y. Bengio · 2016
Cited alongside, same era.
On the training of DNN-based average voice model for speech synthesis
S. Yang, Z. Wu, and L. Xie · 2016
Later among the works it cites.
H. Zen, Y. Agiomyrgiannakis, N. Egberts, F. Henderson, and P. Szczepaniak · 2016
Later among the works it cites.
Deep voice: Real-time neural text-to-speech
S. O. Arik, M. Chrzanowski, A. Coates, G. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, J. Raiman, S. Sengupta, and M. Shoeybi · 2017
Closest in time.
Quasi-recurrent neural networks
J. Bradbury, S. Merity, C. Xiong, and R. Socher · 2017
Closest in time.
C.-C. Hsu, H.-T. Hwang, Y.-C. Wu, Y. Tsao, and H.-M. Wang · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wavenet: A generative model for raw audio
A. v. d. Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Median-based generation of synthetic speech durations using a non-parametric approach
S. Ronanki, O. Watts, S. King, and G. E. Henter · 2016
Cited alongside, same era.
Improved techniques for training gans
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen · 2016
Cited alongside, same era.
C. Li, X. Ma, B. Jiang, X. Li, X. Zhang, X. Liu, Y. Cao, A. Kannan, and Z. Zhu · 2017
Closest in time.
Char2wav: End-to-end speech synthesis
J. Sotelo, S. Mehri, K. Kumar, J. F. Santos, K. Kastner, A. Courville, and Y. Bengio · 2017
Closest in time.
Tacotron: Towards end-to-end speech synthesis
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, et al · 2017
Closest in time.