2021

Guided-TTS: A Diffusion Model for Text-to-Speech via Classifier Guidance

Kim, Heeseung, Kim, Sungwon, Yoon, Sungroh

Understand

We propose Guided-TTS, a high-quality text-to-speech (TTS) model that does not require any transcript of target speaker using classifier guidance.

  • Guided-TTS combines an unconditional diffusion probabilistic model with a separately trained phoneme classifier for classifier guidance.
  • Our unconditional diffusion model learns to generate speech without any context from untranscribed speech data.
  • For TTS synthesis, we guide the generative process of the diffusion model with a phoneme classifier trained on a large-scale speech recognition dataset.

Reading the bibliography…