2018

FloWaveNet : A Generative Flow for Raw Audio

Kim, Sungwon, Lee, Sang-gil, Song, Jongyoon et al.

Understand

Most modern text-to-speech architectures use a WaveNet vocoder for synthesizing high-fidelity waveform audio, but there have been limitations, such as high inference time, in its practical application due to its ancestral sampling scheme.

  • The recently suggested Parallel WaveNet and ClariNet have achieved real-time audio synthesis capability by incorporating inverse autoregressive flow for parallel sampling.
  • However, these approaches require a two-stage training pipeline with a well-trained teacher network and can only produce natural sound by using probability distillation along with auxiliary loss terms.
  • We propose FloWaveNet, a flow-based generative model for raw audio synthesis.

Reading the bibliography…