2021

Alternate Endings: Improving Prosody for Incremental Neural TTS with Predicted Future Text Input

Stephenson, Brooke, Hueber, Thomas, Girin, Laurent et al.

Understand

The prosody of a spoken word is determined by its surrounding context.

  • In incremental text-to-speech synthesis, where the synthesizer produces an output before it has access to the complete input, the full context is often unknown which can result in a loss of naturalness in the synthesized speech.
  • In this paper, we investigate whether the use of predicted future text can attenuate this loss.
  • We compare several test conditions of next future word: (a) unknown (zero-word), (b) language model predicted, (c) randomly predicted and (d) ground-truth.

Reading the bibliography…