2019

Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice Cloning

Zhang, Yu, Weiss, Ron J., Zen, Heiga et al.

Understand

We present a multispeaker, multilingual text-to-speech (TTS) synthesis model based on Tacotron that is able to produce high quality speech in multiple languages.

  • Moreover, the model is able to transfer voices across languages, e.g.
  • synthesize fluent Spanish speech using an English speaker's voice, without training on any bilingual or parallel examples.
  • Such transfer works across distantly related languages, e.g.

Reading the bibliography…