Fetching the paper…
Reading the bibliography…
Existing emotional speech synthesis methods often utilize an utterance-level style embedding extracted from reference audio, neglecting the inherent multi-scale property of speech prosody.
“IEMOCAP: interactive emotional dyadic motion capture database,”
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N. Chang, Sungbok Lee, and Shrikanth S. Narayanan, · 2008
Earlier work this paper cites.
“A kernel two-sample test,”
Arthur Gretton, Karsten M. Borgwardt, Malte J. Rasch, Bernhard Schölkopf, and Alexander J. Smola, · 2012
Earlier work this paper cites.
“The blizzard challenge 2013,”
S. King and Vasilis Karaiskos, · 2013
Earlier work this paper cites.
“Montreal forced aligner: Trainable text-speech alignment using kaldi.,”
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger, · 2017
Earlier work this paper cites.
“3-d convolutional recurrent neural networks with attention model for speech emotion recognition,”
Mingyi Chen, Xuanji He, Jing Yang, and Han Zhang, · 2018
Earlier work this paper cites.
“Generalized end-to-end loss for speaker verification,”
Li Wan, Quan Wang, Alan Papir, and Ignacio Lopez-Moreno, · 2018
Earlier work this paper cites.
“The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,”
Steven R Livingstone and Frank A Russo, · 2018
Earlier work this paper cites.
Adaeze Adigwe, Noé Tits, Kevin El Haddad, Sarah Ostadabbas, and Thierry Dutoit, · 2018
Earlier work this paper cites.
“An open source emotional speech corpus for human robot interaction applications,”
Jesin James, Li Tian, and Catherine Inez Watson, · 2018
Cited alongside, same era.
“Denoising diffusion probabilistic models,”
Jonathan Ho, Ajay Jain, and Pieter Abbeel, · 2020
Cited alongside, same era.
“Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
Cited alongside, same era.
“Score-based generative modeling through stochastic differential equations,”
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole, · 2021
Cited alongside, same era.
“Diffusion models beat gans on image synthesis,”
Prafulla Dhariwal and Alexander Quinn Nichol, · 2021
Cited alongside, same era.
“Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition,”
“Wavlm: Large-scale self-supervised pre-training for full stack speech processing,”
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, Jian Wu, Long Zhou, Shuo Ren, Yanmin Qian, Yao Qian, Jian Wu, Michael Zeng, Xiangzhan Yu, and Furu Wei, · 2022
Later among the works it cites.
“Fine-grained style control in transformer-based text-to-speech synthesis,”
Li-Wei Chen and Alexander Rudnicky, · 2022
Later among the works it cites.
“Emodiff: Intensity controllable emotional text-to-speech with soft-label guidance,”
Yiwei Guo, Chenpeng Du, Xie Chen, and Kai Yu, · 2023
Later among the works it cites.
“Emomix: Emotion mixing via diffusion models for emotional speech synthesis,”
Haobin Tang, Xulong Zhang, Jianzong Wang, Ning Cheng, and Jing Xiao, · 2023
Later among the works it cites.
“Qi-tts: Questioning intonation control for emotional speech synthesis,”
Haobin Tang, Xulong Zhang, Jianzong Wang, Ning Cheng, and Jing Xiao, · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiong Cai, Dongyang Dai, Zhiyong Wu, Xiang Li, Jingbei Li, and Helen Meng, · 2021
Cited alongside, same era.
“Grad-tts: A diffusion probabilistic model for text-to-speech,”
Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, and Mikhail Kudinov, · 2021
Cited alongside, same era.
“Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,”
Kun Zhou, Berrak Sisman, Rui Liu, and Haizhou Li, · 2021
Cited alongside, same era.
Yingzhi Wang, Mirco Ravanelli, Alaa Nfissi, and Alya Yacoubi, · 2023
Later among the works it cites.
“Improved cross-corpus speech emotion recognition using deep local domain adaptation,”
Zhao Huijuan, YE Ning, and Wang Ruchuan, · 2023
Later among the works it cites.