Fetching the paper…
Reading the bibliography…
We introduce Matcha-TTS, a new encoder-decoder architecture for speedy TTS acoustic modelling, trained using optimal-transport conditional flow matching (OT-CFM).
K. Prahallad, A. Vadapalli, N. Elluru, G. Mantena, B. Pulugundla et al. , “The Blizzard Challenge 2013 – Indian language task,” in Proc. Blizzard Challenge Workshop , 2013
2013
Earlier work this paper cites.
R. T. Q. Chen, Y. Rubanova, J. Bettencourt et al. , “Neural ordinary differential equations,” in Proc. NeurIPS , 2018
2018
Earlier work this paper cites.
P. Shaw, J. Uszkoreit, and A. Vaswani, “Self-attention with relative position representations,” in Proc. NAACL , 2018
2018
Earlier work this paper cites.
Y. Song and S. Ermon, “Generative modeling by estimating gradients of the data distribution,” Proc. NeurIPS , 2019
2019
Earlier work this paper cites.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “FastSpeech: Fast, robust and controllable text to speech,” in Proc. NeurIPS , 2019
2019
Earlier work this paper cites.
O. Watts, G. E. Henter, J. Fong, and C. Valentini-Botinhao, “Where do the improvements come from in sequence-to-sequence neural TTS?” in Proc. SSW , 2019, pp. 217–222
2019
Earlier work this paper cites.
R. Prenger, R. Valle, and B. Catanzaro, “WaveGlow: A flow-based generative network for speech synthesis,” in Proc. ICASSP , 2019, pp. 3617–3621
2019
Earlier work this paper cites.
É. Székely, G. E. Henter, J. Beskow, and J. Gustafson, “Spontaneous conversational speech synthesis from found data,” in Proc. Interspeech , 2019, pp. 4435–4439
2019
Earlier work this paper cites.
J. Kim, S. Kim, J. Kong, and S. Yoon, “Glow-TTS: A generative flow for text-to-speech via monotonic alignment search,” in Proc. NeurIPS , 2020, pp. 8067–8077
2020
Earlier work this paper cites.
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” in Proc. NeurIPS , 2020, pp. 17 022–17 033
2020
Earlier work this paper cites.
P. Dhariwal and A. Nichol, “Diffusion models beat GANs on image synthesis,” in Proc. NeurIPS , 2021, pp. 8780–8794
2021
Earlier work this paper cites.
N. Chen, Y. Zhang, H. Zen, R. J. Weiss, M. Norouzi, and W. Chan, “WaveGrad: Estimating gradients for waveform generation,” in Proc. ICLR , 2021
2021
Earlier work this paper cites.
N. Chen, Y. Zhang, H. Zen, R. J. Weiss, M. Norouzi, N. Dehak, and W. Chan, “WaveGrad 2: Iterative refinement for text-to-speech synthesis,” in Proc. Interspeech , 2021, pp. 3765–3769
2021
Earlier work this paper cites.
V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, and M. Kudinov, “Grad-TTS: A diffusion probabilistic model for text-to-speech,” in Proc. ICML , 2021, pp. 8599–8608
2021
Earlier work this paper cites.
M. Jeong, H. Kim, S. J. Cheon, B. J. Choi, and N. S. Kim, “Diff-TTS: A denoising diffusion model for text-to-speech,” in Proc. Interspeech , 2021, pp. 3605–3609
2021
Cited alongside, same era.
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, “DiffWave: A versatile diffusion model for audio synthesis,” in Proc. ICLR , 2021
2021
Cited alongside, same era.
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” in Proc. ICLR , 2021
2021
Cited alongside, same era.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “FastSpeech 2: Fast and high-quality end-to-end text to speech,” in Proc. ICLR , 2021
2021
Cited alongside, same era.
J. Kim, J. Kong, and J. Son, “VITS: Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,” in Proc. ICML , 2021, pp. 5530–5540
R. Huang, Z. Zhao, H. Liu, J. Liu, C. Cui, and Y. Ren, “ProDiff: Progressive fast diffusion model for high-quality text-to-speech,” in Proc. MM , 2022, pp. 2595–2605
2022
Later among the works it cites.
S. Mehta, É. Székely, J. Beskow, and G. E. Henter, “Neural HMMs are all you need (for high-quality attention-free TTS),” in Proc. ICASSP , 2022, pp. 7457–7461
2022
Later among the works it cites.
S. Alexanderson, R. Nagy, J. Beskow, and G. E. Henter, “Listen, denoise, action! Audio-driven motion synthesis with diffusion models,” ACM ToG , vol. 42, no. 4, 2023, article 44
2023
Closest in time.
S. Mehta, S. Wang, S. Alexanderson, J. Beskow, É. Székely, and G. E. Henter, “Diff-TTSG: Denoising probabilistic integrated speech and gesture synthesis,” in Proc. SSW , 2023
2023
Closest in time.
Y. Lipman, R. T. Q. Chen, H. Ben-Hamu et al. , “Flow matching for generative modeling,” in Proc. ICLR , 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
H. Zhang, Z. Huang, Z. Shang, P. Zhang, and Y. Yan, “LinearSpeech: Parallel text-to-speech with linear complexity,” in Proc. Interspeech , 2021, pp. 4129–4133
2021
Cited alongside, same era.
2021
Cited alongside, same era.
U. Wennberg and G. E. Henter, “The case for translation-invariant self-attention in Transformer-based language models,” in Proc. ACL-IJCNLP Vol. 2 , 2021, pp. 130–140
2021
Cited alongside, same era.
M. Bernard and H. Titeux, “Phonemizer: Text to phones transcription for multiple languages in Python,” J. Open Source Softw. , vol. 6, no. 68, p. 3958, 2021
2021
Cited alongside, same era.
J. Taylor and K. Richmond, “Confidence intervals for ASR-based TTS evaluation,” in Proc. Interspeech , 2021
2021
Cited alongside, same era.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in Proc. CVPR , 2022, pp. 10 684–10 695
2022
Cited alongside, same era.
M. S. Albergo and E. Vanden-Eijnden, “Building normalizing flows with stochastic interpolants,” in Proc. ICLR , 2022
2022
Cited alongside, same era.
2023
Closest in time.
J. Betker, “Better speech synthesis through scaling,” arXiv preprint arXiv:2305.07243 , 2023
2023
Closest in time.
S. Mehta, A. Kirkland, H. Lameris, J. Beskow, É. Székely, and G. E. Henter, “OverFlow: Putting flows on top of neural transducers for better TTS,” in Proc. Interspeech , 2023
2023
Closest in time.
S.-g. Lee, W. Ping, B. Ginsburg, B. Catanzaro, and S. Yoon, “BigVGAN: A universal neural vocoder with large-scale training,” in Proc. ICLR , 2023
2023
Closest in time.
X. Liu et al. , “Flow straight and fast: Learning to generate and transfer data with rectified flow,” in Proc. ICLR , 2023
2023
Closest in time.
Z. Ye, W. Xue, X. Tan, J. Chen, Q. Liu, and Y. Guo, “CoMoSpeech: One-step speech and singing voice synthesis via consistency model,” in Proc. ACM MM , 2023, pp. 1831–1839
2023
Closest in time.
2023
Closest in time.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. Mcleavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in Proc. ICML , 2023, pp. 28 492–28 518
2023
Closest in time.
C.-H. Chiang, W.-P. Huang, and H. yi Lee, “Why we should report the details in subjective evaluation of TTS more rigorously,” in Proc. Interspeech , 2023, pp. 5551–5555
2023
Closest in time.
A. Kirkland, S. Mehta, H. Lameris, G. E. Henter, E. Szekely et al. , “Stuck in the MOS pit: A critical analysis of MOS test methodology in TTS evaluation,” in Proc. SSW , 2023
2023
Closest in time.