Fetching the paper…
Reading the bibliography…
We introduce an approach to multilingual speech synthesis which uses the meta-learning concept of contextual parameter generation and produces natural-sounding multilingual speech using more languages and less training data than previous approaches.
R. W. Soukoreff and I. S. MacKenzie, “Measuring errors in text entry tasks: an application of the Levenshtein string distance statistic,” in
2001
Earlier work this paper cites.
T. Kudo, “MeCab: Yet Another Part-of-Speech and Morphological Analyzer,”
2013
Earlier work this paper cites.
M. Yao, “A Romaji/Kana conversion library for Python,”
2015
Earlier work this paper cites.
“Method for the subjective assessment of intermediate quality levels of coding systems,” International Telecommunication Union, Geneva, Recommendation BS.1534, Oct. 2015
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
Y. Ganin, E. Ustinova, H. Ajakan, P. Germain, H. Larochelle, F. Laviolette, M. Marchand, and V. Lempitsky, “Domain-Adversarial Training of Neural Networks,”
2016
Earlier work this paper cites.
L. Yu, “Pinyin,”
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. Shen, R. Pang, R. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. J. Skerrv-Ryan, R. Saurous, Y. Agiomvrgiannakis, and Y. Wu, “Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions,” in
2018
Earlier work this paper cites.
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. van den Oord, S. Dieleman, and K. Kavukcuoglu, “Efficient Neural Audio Synthesis,” in
2018
Earlier work this paper cites.
Y. Wang, D. Stanton, Y. Zhang, R.-S. Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, and R. A. Saurous, “Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,” in
2018
Cited alongside, same era.
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, Z. Chen, P. Nguyen, R. Pang, I. L. Moreno, and Y. Wu, “Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis,” in
2018
Cited alongside, same era.
E. A. Platanios, M. Sachan, G. Neubig, and T. M. Mitchell, “Contextual Parameter Generation for Universal Neural Machine Translation,” in
2018
Cited alongside, same era.
D. Sachan and G. Neubig, “Parameter Sharing Methods for Multilingual Self-Attentional Translation Models,” in
2018
Cited alongside, same era.
A. Prakash, A. Leela Thomas, S. Umesh, and H. A Murthy, “Building Multilingual End-to-End Speech Synthesisers for Indian Languages,” in
2019
Later among the works it cites.
E. Nachmani and L. Wolf, “Unsupervised Polyglot Text-to-speech,” in
2019
Later among the works it cites.
M. Chen, M. Chen, S. Liang, J. Ma, L. Chen, S. Wang, and J. Xiao, “Cross-Lingual, Multi-Speaker Text-To-Speech Synthesis Using Neural Speaker Embedding,” in
2019
Later among the works it cites.
Y. Cao, X. Wu, S. Liu, J. Yu, X. Li, Z. Wu, X. Liu, and H. M. Meng, “End-to-end Code-switched TTS with Mix of Monolingual Recordings,”
2019
Later among the works it cites.
Fatchord, “WaveRNN Vocoder + TTS,”
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
Y. Taigman, L. Wolf, A. Polyak, and E. Nachmani, “VoiceLoop: Voice Fitting and Synthesis via a Phonological Loop,” in
2018
Cited alongside, same era.
H. Tachibana, K. Uenoyama, and S. Aihara, “Efficiently Trainable Text-to-Speech System Based on Deep Convolutional Networks with Guided Attention,”
2018
Cited alongside, same era.
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brébisson, Y. Bengio, and A. C. Courville, “MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis,” in
2019
Cited alongside, same era.
W. Hsu, Y. Zhang, R. J. Weiss, H. Zen, Y. Wu, Y. Cao, and Y. Wang, “Hierarchical Generative Modeling for Controllable Speech Synthesis,” in
2019
Cited alongside, same era.
2019
Later among the works it cites.
R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber, “Common Voice: A Massively-Multilingual Speech Corpus,” in
2020
Closest in time.
Y.-J. Chen, T. Tu, C. chieh Yeh, and H.-Y. Lee, “End-to-End Text-to-Speech for Low-Resource Languages by Cross-Lingual Transfer Learning,” in
2079
Closest in time.
Y. Zhang, R. Weiss, H. Zen, Y. Wu, Z. Chen, R. Skerry-Ryan, Y. Jia, A. Rosenberg, and B. Ramabhadran, “Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice Cloning,” in
2084
Closest in time.