2021

Improving Performance of Seen and Unseen Speech Style Transfer in End-to-end Neural TTS

An, Xiaochun, Soong, Frank K., Xie, Lei

Understand

End-to-end neural TTS training has shown improved performance in speech style transfer.

  • However, the improvement is still limited by the training data in both target styles and speakers.
  • Inadequate style transfer performance occurs when the trained TTS tries to transfer the speech to a target style from a new speaker with an unknown, arbitrary style.
  • In this paper, we propose a new approach to style transfer for both seen and unseen styles, with disjoint, multi-style datasets, i.e., datasets of different styles are recorded, each individual style is by one speaker with multiple utterances.

Reading the bibliography…