Fetching the paper…
Reading the bibliography…
Although current neural text-to-speech (TTS) models are able to generate high-quality speech, intensity controllable emotional TTS is still a challenging task.
D. Parikh and K. Grauman, “Relative attributes,” in Proc. ICCV , 2011, pp. 503–510
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz et al. , “The kaldi speech recognition toolkit,” in IEEE 2011 workshop on automatic speech recognition and understanding , no. CONF. IEEE Signal Processing Society, 2011
2011
Earlier work this paper cites.
2017
Earlier work this paper cites.
X. Zhu, S. Yang, G. Yang, and L. Xie, “Controlling emotion strength with relative attribute for end-to-end speech synthesis,” in Proc. IEEE ASRU , 2019, pp. 192–199
2019
Earlier work this paper cites.
Y. Song and S. Ermon, “Generative modeling by estimating gradients of the data distribution,” Proc. NeurIPS , vol. 32, 2019
2019
Earlier work this paper cites.
G. Sun, Y. Zhang, R. J. Weiss, Y. Cao, H. Zen, and Y. Wu, “Fully-hierarchical fine-grained prosody modeling for interpretable speech synthesis,” in Proc. IEEE ICASSP , 2020, pp. 6264–6268
2020
Earlier work this paper cites.
S.-Y. Um, S. Oh, K. Byun, I. Jang, C. Ahn, and H.-G. Kang, “Emotional speech synthesis with rich and granularized control,” in Proc. IEEE ICASSP , 2020, pp. 7254–7258
2020
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Proc. NeurIPS , vol. 33, pp. 6840–6851, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, and M. Kudinov, “Grad-tts: A diffusion probabilistic model for text-to-speech,” in Proc. ICML , 2021, pp. 8599–8608
2021
Earlier work this paper cites.
J. Kim, J. Kong, and J. Son, “Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,” in Proc. ICML , vol. 139, 2021, pp. 5530–5540
2021
Earlier work this paper cites.
C. Du and K. Yu, “Phone-level prosody modelling with GMM-based MDN for diverse and controllable speech synthesis,” IEEE/ACM Trans. ASLP. , vol. 30, pp. 190–201, 2021
2021
Cited alongside, same era.
M. Kim, S. J. Cheon, B. J. Choi, J. J. Kim, and N. S. Kim, “Expressive text-to-speech using style tag,” in Proc. ISCA Interspeech , 2021, pp. 4663–4667
2021
Cited alongside, same era.
Y. Lei, S. Yang, and L. Xie, “Fine-grained emotion strength transfer, control and prediction for emotional speech synthesis,” in Proc. IEEE SLT , 2021, pp. 423–430
2021
Cited alongside, same era.
B. Schnell and P. N. Garner, “Improving emotional tts with an emotion intensity input from unsupervised extraction,” in Proc. 11th ISCA Speech Synthesis Workshop (SSW 11) , 2021, pp. 60–65
2021
Cited alongside, same era.
H. Choi and M. Hahn, “Sequence-to-sequence emotional voice conversion with strength control,” IEEE Access , vol. 9, pp. 42 674–42 687, 2021
K. Zhou, B. Sisman, R. Liu, and H. Li, “Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,” in Proc. IEEE ICASSP , 2021, pp. 920–924
2021
Later among the works it cites.
C. Du, Y. Guo, X. Chen, and K. Yu, “VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature,” in Proc. ISCA Interspeech , 2022, pp. 1596–1600
2022
Closest in time.
Y. Guo, C. Du, and K. Yu, “Unsupervised word-level prosody tagging for controllable speech synthesis,” in Proc. IEEE ICASSP , 2022, pp. 7597–7601
2022
Closest in time.
Y. Lei, S. Yang, X. Wang, and L. Xie, “Msemotts: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis,” IEEE/ACM Trans. ASLP. , vol. 30, pp. 853–864, 2022
2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
P. Dhariwal and A. Nichol, “Diffusion models beat gans on image synthesis,” Proc. NeurIPS , vol. 34, pp. 8780–8794, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” in Proc. ICLR , 2021
2021
Cited alongside, same era.
M. Jeong, H. Kim, S. J. Cheon, B. J. Choi, and N. S. Kim, “Diff-TTS: A Denoising Diffusion Model for Text-to-Speech,” in Proc. ISCA Interspeech , 2021, pp. 3605–3609
2021
Cited alongside, same era.
N. Chen, Y. Zhang, H. Zen, R. J. Weiss, M. Norouzi, and W. Chan, “Wavegrad: Estimating gradients for waveform generation,” in Proc. ICLR , 2021
2021
Cited alongside, same era.
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, “Diffwave: A versatile diffusion model for audio synthesis,” in Proc. ICLR , 2021
2021
Cited alongside, same era.
2022
Closest in time.
C.-B. Im, S.-H. Lee, S.-B. Kim, and S.-W. Lee, “Emoq-tts: Emotion intensity quantization for fine-grained controllable emotional text-to-speech,” in Proc. IEEE ICASSP , 2022, pp. 6317–6321
2022
Closest in time.
K. Zhou, B. Sisman, R. Rana, B. W. Schuller, and H. Li, “Emotion intensity and its control for emotional voice conversion,” IEEE Transactions on Affective Computing , 2022
2022
Closest in time.
J. Liu, C. Li, Y. Ren, F. Chen, and Z. Zhao, “Diffsinger: Singing voice synthesis via shallow diffusion mechanism,” in Proc. AAAI , 2022, pp. 11 020–11 028
2022
Closest in time.
R. Huang, M. W. Y. Lam, J. Wang, D. Su, D. Yu, Y. Ren, and Z. Zhao, “FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis,” in Proc. IJCAI , 2022, pp. 4157–4163
2022
Closest in time.
M. W. Y. Lam, J. Wang, D. Su, and D. Yu, “BDDM: Bilateral denoising diffusion models for fast and high-quality speech synthesis,” in Proc. ICLR , 2022
2022
Closest in time.
H. Kim, S. Kim, and S. Yoon, “Guided-tts: A diffusion model for text-to-speech via classifier guidance,” in Proc. ICML , 2022, pp. 11 119–11 133
2022
Closest in time.