Fetching the paper…
Reading the bibliography…
Although diffusion models in text-to-speech have become a popular choice due to their strong generative ability, the intrinsic complexity of sampling from diffusion models harms their efficiency.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz et al. , “The Kaldi speech recognition toolkit,” in Proc. IEEE ASRU . IEEE Signal Processing Society, 2011
2011
Earlier work this paper cites.
K. Ito and L. Johnson, “The LJ speech dataset,” 2017
2017
Earlier work this paper cites.
H. Zen, R. Clark, R. J. Weiss, V. Dang, Y. Jia, Y. Wu, Y. Zhang, and Z. Chen, “LibriTTS: A corpus derived from librispeech for text-to-speech,” in Interspeech , 2019
2019
Earlier work this paper cites.
C.-C. Lo, S.-W. Fu, W.-C. Huang, X. Wang, J. Yamagishi, Y. Tsao, and H.-M. Wang, “MOSNet: Deep learning based objective assessment for voice conversion,” in Proc. Interspeech 2019 , 2019
2019
Earlier work this paper cites.
M. Bińkowski, J. Donahue, S. Dieleman, A. Clark, E. Elsen, N. Casagrande, L. C. Cobo, and K. Simonyan, “High fidelity speech synthesis with adversarial networks,” in Proc. ICLR , 2020
2020
Earlier work this paper cites.
J. Kim, S. Kim, J. Kong, and S. Yoon, “Glow-TTS: A generative flow for text-to-speech via monotonic alignment search,” Proc. NeurIPS , vol. 33, pp. 8067–8077, 2020
2020
Earlier work this paper cites.
J. Kong, J. Kim, and J. Bae, “Hifi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” Proc. NeurIPS , vol. 33, pp. 17 022–17 033, 2020
2020
Earlier work this paper cites.
J. Kim, J. Kong, and J. Son, “Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,” in Proc. ICML . PMLR, 2021, pp. 5530–5540
2021
Earlier work this paper cites.
R. Valle, K. J. Shih, R. Prenger, and B. Catanzaro, “Flowtron: an autoregressive flow-based generative network for text-to-speech synthesis,” in Proc. ICLR , 2021
2021
Earlier work this paper cites.
V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, and M. Kudinov, “Grad-TTS: A diffusion probabilistic model for text-to-speech,” in Proc. ICML . PMLR, 2021, pp. 8599–8608
2021
Earlier work this paper cites.
P. Dhariwal and A. Nichol, “Diffusion models beat gans on image synthesis,” Proc. NeurIPS , vol. 34, pp. 8780–8794, 2021
2021
Earlier work this paper cites.
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in Proc. ICLR , 2021
2021
Earlier work this paper cites.
C. Du, Y. Guo, X. Chen, and K. Yu, “VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature,” in Proc. Interspeech 2022 , pp. 1596–1600
2022
Cited alongside, same era.
J. Liu, C. Li, Y. Ren, F. Chen, and Z. Zhao, “DiffSinger: Singing voice synthesis via shallow diffusion mechanism,” in Proceedings of the AAAI conference on artificial intelligence , vol. 36, no. 10, 2022, pp. 11 020–11 028
2022
Cited alongside, same era.
H. Kim, S. Kim, and S. Yoon, “Guided-TTS: A diffusion model for text-to-speech via classifier guidance,” in Proc. ICML . PMLR, 2022, pp. 11 119–11 133
2022
Cited alongside, same era.
J. Tae, H. Kim, and T. Kim, “EdiTTS: Score-based Editing for Controllable Text-to-Speech,” in Proc. Interspeech 2022 , 2022, pp. 421–425
2022
Cited alongside, same era.
X. Liu, C. Gong, and Q. Liu, “Flow straight and fast: Learning to generate and transfer data with rectified flow,” in NeurIPS 2022 Workshop on Score-Based Methods , 2022
2022
Later among the works it cites.
C. Du, Y. Guo, X. Chen, and K. Yu, “Speaker adaptive text-to-speech with timbre-normalized vector-quantized feature,” IEEE/ACM Trans. ASLP. , 2023
2023
Closest in time.
Z. Liu, Y. Guo, and K. Yu, “DiffVoice: Text-to-speech with latent diffusion,” in Proc. IEEE ICASSP , 2023
2023
Closest in time.
2023
Closest in time.
Y. Guo, C. Du, X. Chen, and K. Yu, “EmoDiff: Intensity controllable emotional text-to-speech with soft-label guidance,” in Proc. IEEE ICASSP , 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, M. S. Kudinov, and J. Wei, “Diffusion-based voice conversion with fast maximum likelihood sampling scheme,” in Proc. ICLR , 2022
2022
Cited alongside, same era.
T. Salimans and J. Ho, “Progressive distillation for fast sampling of diffusion models,” in Proc. ICLR , 2022
2022
Cited alongside, same era.
Z. Xiao, K. Kreis, and A. Vahdat, “Tackling the generative learning trilemma with denoising diffusion GANs,” in Proc. ICLR , 2022
2022
Cited alongside, same era.
R. Huang, M. W. Y. Lam, J. Wang, D. Su, D. Yu, Y. Ren, and Z. Zhao, “FastDiff: A fast conditional diffusion model for high-quality speech synthesis,” in Proc. IJCAI , 2022, pp. 4157–4163
2022
Cited alongside, same era.
M. W. Y. Lam, J. Wang, D. Su, and D. Yu, “BDDM: Bilateral denoising diffusion models for fast and high-quality speech synthesis,” in Proc. ICLR , 2022
2022
Cited alongside, same era.
R. Huang, Z. Zhao, H. Liu, J. Liu, C. Cui, and Y. Ren, “ProDiff: Progressive fast diffusion model for high-quality text-to-speech,” in Proceedings of the 30th ACM International Conference on Multimedia , 2022, pp. 2595–2605
2022
Cited alongside, same era.
C. Lu, Y. Zhou, F. Bao, J. Chen, C. Li, and J. Zhu, “DPM-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps,” Proc. NeurIPS , vol. 35, pp. 5775–5787, 2022
2022
Cited alongside, same era.
2023
Closest in time.
J. Chen, X. Song, Z. Peng, B. Zhang, F. Pan, and Z. Wu, “LightGrad: Lightweight diffusion probabilistic model for text-to-speech,” in Proc. IEEE ICASSP , 2023
2023
Closest in time.
2023
Closest in time.
Y. Song, P. Dhariwal, M. Chen, and I. Sutskever, “Consistency models,” in Proc. ICML . PMLR, 2023, pp. 32 211–32 252
2023
Closest in time.
Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le, “Flow matching for generative modeling,” in Proc. ICLR , 2023
2023
Closest in time.
A. Tong, N. Malkin, G. Huguet, Y. Zhang, J. Rector-Brooks, K. FATRAS, G. Wolf, and Y. Bengio, “Improving and generalizing flow-based generative models with minibatch optimal transport,” in ICML Workshop on New Frontiers in Learning, Control, and Dynamical Systems , 2023
2023
Closest in time.
2023
Closest in time.