Fetching the paper…

DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors · Around