2018

Deep Encoder-Decoder Models for Unsupervised Learning of Controllable Speech Synthesis

Henter, Gustav Eje, Lorenzo-Trueba, Jaime, Wang, Xin et al.

Understand

Generating versatile and appropriate synthetic speech requires control over the output expression separate from the spoken text.

  • Important non-textual speech variation is seldom annotated, in which case output control must be learned in an unsupervised fashion.
  • In this paper, we perform an in-depth study of methods for unsupervised learning of control in statistical speech synthesis.
  • For example, we show that popular unsupervised training heuristics can be interpreted as variational inference in certain autoencoder models.

Reading the bibliography…