Fetching the paper…
Reading the bibliography…
There has been significant progress in emotional Text-To-Speech (TTS) synthesis technology in recent years.
R. Kubichek, “Mel-cepstral distance measure for objective speech quality assessment,” in Proceedings of IEEE Pacific Rim Conference on Communications Computers and Signal Processing , vol. 1. IEEE, 1993, pp. 125–128
1993
Earlier work this paper cites.
R. Plutchik, “The nature of emotions: Human emotions have deep evolutionary roots, a fact that may explain their complexity and provide tools for clinical practice,” American Scientist , vol. 89, no. 4, pp. 344–350, 2001
2001
Earlier work this paper cites.
C. Busso, M. Bulut, C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “IEMOCAP: interactive emotional dyadic motion capture database,” Lang. Resour. Evaluation , vol. 42, no. 4, pp. 335–359, 2008
2008
Earlier work this paper cites.
D. Parikh and K. Grauman, “Relative attributes,” in 2011 International Conference on Computer Vision . IEEE, 2011, pp. 503–510
2011
Earlier work this paper cites.
R. Plutchik and H. Kellerman, Theories of emotion . Academic Press, 2013, vol. 1
2013
Earlier work this paper cites.
A. Braniecka, E. Trzebińska, A. Dowgiert, and A. Wytykowska, “Mixed emotions and coping: The benefits of secondary emotions,” PloS one , vol. 9, no. 8, p. e103940, 2014
2014
Earlier work this paper cites.
J. Johnson, A. Alahi, and L. Fei-Fei, “Perceptual losses for real-time style transfer and super-resolution,” in European Conference on Computer Vision (ECCV) Part II 14 , 2016, pp. 694–711
2016
Earlier work this paper cites.
M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, “Montreal forced aligner: Trainable text-speech alignment using kaldi.” in Interspeech , vol. 2017, 2017, pp. 498–502
2017
Earlier work this paper cites.
Y. Wang, D. Stanton, Y. Zhang, R.-S. Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, and R. A. Saurous, “Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,” in International Conference on Machine Learning . PMLR, 2018, pp. 5180–5189
2018
Earlier work this paper cites.
M. Chen, X. He, J. Yang, and H. Zhang, “3-d convolutional recurrent neural networks with attention model for speech emotion recognition,” IEEE Signal Processing Letters , vol. 25, no. 10, pp. 1440–1444, 2018
2018
Earlier work this paper cites.
S. Ma, D. Mcduff, and Y. Song, “Neural tts stylization with adversarial and collaborative games,” in International Conference on Learning Representations , 2019
2019
Earlier work this paper cites.
X. Zhu, S. Yang, G. Yang, and L. Xie, “Controlling emotion strength with relative attribute for end-to-end speech synthesis,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 192–199
2019
Earlier work this paper cites.
S.-Y. Um, S. Oh, K. Byun, I. Jang, C. Ahn, and H.-G. Kang, “Emotional speech synthesis with rich and granularized control,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7254–7258
2020
Cited alongside, same era.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in Neural Information Processing Systems , vol. 33, pp. 6840–6851, 2020
2020
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in Neural Information Processing Systems , vol. 33, pp. 12 449–12 460, 2020
2020
Cited alongside, same era.
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 17 022–17 033
2020
Cited alongside, same era.
X. Cai, D. Dai, Z. Wu, X. Li, J. Li, and H. Meng, “Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 5734–5738
2021
Later among the works it cites.
R. Huang, Y. Ren, J. Liu, C. Cui, and Z. Zhao, “Generspeech: Towards style transfer for generalizable out-of-domain text-to-speech,” in Advances in Neural Information Processing Systems , vol. 35, 2022, pp. 10 970–10 983
2022
Later among the works it cites.
T. Li, X. Wang, Q. Xie, Z. Wang, M. Jiang, and L. Xie, “Cross-speaker Emotion Transfer Based On Prosody Compensation for End-to-End Speech Synthesis,” in Proc. Interspeech 2022 , 2022, pp. 5498–5502
2022
Later among the works it cites.
C.-B. Im, S.-H. Lee, S.-B. Kim, and S.-W. Lee, “Emoq-tts: Emotion intensity quantization for fine-grained controllable emotional text-to-speech,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 6317–6321
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Liu, B. Sisman, G. Gao, and H. Li, “Expressive tts training with frame and style reconstruction loss,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 1806–1818, 2021
2021
Cited alongside, same era.
H. Choi and M. Hahn, “Sequence-to-sequence emotional voice conversion with strength control,” IEEE Access , vol. 9, pp. 42 674–42 687, 2021
2021
Cited alongside, same era.
P. Dhariwal and A. Q. Nichol, “Diffusion models beat gans on image synthesis,” in Advances in Neural Information Processing Systems , vol. 34, 2021, pp. 8780–8794
2021
Cited alongside, same era.
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” in International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, and M. Kudinov, “Grad-tts: A diffusion probabilistic model for text-to-speech,” in International Conference on Machine Learning . PMLR, 2021, pp. 8599–8608
2021
Cited alongside, same era.
K. Zhou, B. Sisman, R. Liu, and H. Li, “Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 920–924
2021
Cited alongside, same era.
T. Li, S. Yang, L. Xue, and L. Xie, “Controllable emotion transfer for end-to-end speech synthesis,” in 2021 12th International Symposium on Chinese Spoken Language Processing (ISCSLP) . IEEE, 2021, pp. 1–5
2021
Cited alongside, same era.
2022
Later among the works it cites.
Y. Lei, S. Yang, X. Wang, and L. Xie, “Msemotts: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 853–864, 2022
2022
Later among the works it cites.
K. Zhou, B. Sisman, R. Rana, B. W. Schuller, and H. Li, “Emotion intensity and its control for emotional voice conversion,” IEEE Transactions on Affective Computing , no. 01, pp. 1–1, 2022
2022
Later among the works it cites.
K. Zhou, B. Sisman, R. Rana, B. W. Schuller, and H. Li, “Speech synthesis with mixed emotions,” IEEE Transactions on Affective Computing , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
G. Kim, T. Kwon, and J. C. Ye, “Diffusionclip: Text-guided diffusion models for robust image manipulation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 2426–2435
2022
Later among the works it cites.
R. Huang, Z. Zhao, H. Liu, J. Liu, C. Cui, and Y. Ren, “Prodiff: Progressive fast diffusion model for high-quality text-to-speech,” in Proceedings of the 30th ACM International Conference on Multimedia , 2022, pp. 2595–2605
2022
Later among the works it cites.
Y. Guo, C. Du, X. Chen, and K. Yu, “Emodiff: Intensity controllable emotional text-to-speech with soft-label guidance,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2023, pp. 1–5
2023
Closest in time.