Fetching the paper…
Reading the bibliography…
Modeling voices for multiple speakers and multiple languages in one text-to-speech system has been a challenge for a long time.
H. of the International Phonetic Association
1999
Earlier work this paper cites.
A. B. Bernardo, “Bilingual code-switching as a resource for learning and teaching: Alternative reflections on the language and education issue in the philippines,”
2005
Earlier work this paper cites.
D.-C. Lyu and R.-Y. Lyu, “Language identification on code-switching utterances using multiple cues,” in
2008
Earlier work this paper cites.
D.-C. Lyu, T.-P. Tan, E. S. Chng, and H. Li, “Seame: a mandarin-english code-switching speech corpus in south-east asia,” in
2010
Earlier work this paper cites.
H.-P. Shen, C.-H. Wu, Y.-T. Yang, and C.-S. Hsu, “Cecos: A chinese-english code-switching speech database,” in
2011
Earlier work this paper cites.
B. H. Ahmed and T.-P. Tan, “Automatic speech recognition of code switching speech using 1-best rescoring,” in
2012
Earlier work this paper cites.
N. T. Vu, D.-C. Lyu, J. Weiner, D. Telaar, T. Schlippe, F. Blaicher, E.-S. Chng, T. Schultz, and H. Li, “A first speech recognition system for mandarin-english code-switch conversational speech,” in
2012
Earlier work this paper cites.
D.-C. Lyu, E.-S. Chng, and H. Li, “Language diarization for code-switch conversational speech,” in
2013
Earlier work this paper cites.
M. J. Gales, K. M. Knill, and A. Ragni, “Unicode-based graphemic systems for limited resource languages,” in
2015
Cited alongside, same era.
H. Ming, Y. Lu, Z. Zhang, and M. Dong, “A light-weight method of building an LSTM-RNN-based bilingual TTS system,” in
2017
Cited alongside, same era.
K. Ito, “The LJ Speech Dataset,” https://keithito.com/LJ-Speech-Dataset/, 2017
2017
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan
2018
Cited alongside, same era.
Y. Wang, D. Stanton, Y. Zhang, R.-S. Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, and R. A. Saurous, “Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,” in
2018
Cited alongside, same era.
B. Li, Y. Zhang, T. Sainath, Y. Wu, and W. Chan, “Bytes are all you need: End-to-end multilingual speech recognition and synthesis with bytes,” in
2019
Later among the works it cites.
M. Chen, M. Chen, S. Liang, J. Ma, L. Chen, S. Wang, and J. Xiao, “Cross-lingual, multi-speaker text-to-speech synthesis using neural speaker embedding,”
2019
Later among the works it cites.
2019
Later among the works it cites.
X. Zhou, X. Tian, G. Lee, R. K. Das, and H. Li, “End-to-end code-switching tts with cross-lingual language model,” in
2020
Closest in time.
E. Cooper, C. Lai, Y. Yasuda, F. Fang, X. Wang, N. Chen, and J. Yamagishi, “Zero-shot multi-speaker text-to-speech with state-of-the-art neural speaker embeddings,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, I. L. Moreno, Y. Wu
2018
Cited alongside, same era.
2018
Cited alongside, same era.
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. van den Oord, S. Dieleman, and K. Kavukcuoglu, “Efficient neural audio synthesis,” in
2018
Cited alongside, same era.
B. Li and H. Zen, “multi-language multi-speaker acoustic modeling for lstm-rnn based statistical parametric speech synthesis.”
Cited in the paper.
“The Carnegie Mellon Pronouncing Dictionary,” http://www.speech.cs.cmu.edu/cgi-bin/cmudict
Cited in the paper.
“Mandarin Pinyin to CMU Dictionary Phoneme Set,” https://github.com/kaldi-asr/kaldi/blob/master/egs/hkust/s5/conf/pinyin2cmu
Cited in the paper.
2020
Closest in time.
W. Cai, J. Chen, J. Zhang, and M. Li, “On-the-fly data loader and utterance-level aggregation for speaker and language recognition,”
2020
Closest in time.
Y. Zhang, R. J. Weiss, H. Zen, Y. Wu, Z. Chen, R. Skerry-Ryan, Y. Jia, A. Rosenberg, and B. Ramabhadran, “Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice Cloning,” in
2084
Closest in time.