Understand
We present a lightweight adaptable neural TTS system with high quality output.
- The system is composed of three separate neural network blocks: prosody prediction, acoustic feature prediction and Linear Prediction Coding Net as a neural vocoder.
- This system can synthesize speech with close to natural quality while running 3 times faster than real-time on a standard CPU.
- The modular setup of the system allows for simple adaptation to new voices with a small amount of data.
Built on
Nothing clear enough to list yet.
Similar
Nothing clear enough to list yet.
Then
Nothing clear enough to list yet.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…