2020

Multi-band MelGAN: Faster Waveform Generation for High-Quality Text-to-Speech

Yang, Geng, Yang, Shan, Liu, Kai et al.

Understand

In this paper, we propose multi-band MelGAN, a much faster waveform generation model targeting to high-quality text-to-speech.

  • Specifically, we improve the original MelGAN by the following aspects.
  • First, we increase the receptive field of the generator, which is proven to be beneficial to speech generation.
  • Second, we substitute the feature matching loss with the multi-resolution STFT loss to better measure the difference between fake and real speech.

Reading the bibliography…