Fetching the paper…
Reading the bibliography…
Emotional voice conversion aims to transform emotional prosody in speech while preserving the linguistic content and speaker identity.
“Emotion and personality.,”
Magda B Arnold, · 1960
Earlier work this paper cites.
“A circumplex model of affect.,”
James A Russell, · 1980
Earlier work this paper cites.
“An argument for basic emotions,”
Paul Ekman, · 1992
Earlier work this paper cites.
“Pragmatics and intonation,”
Julia Hirschberg, · 2004
Earlier work this paper cites.
“Prosody conversion from neutral speech to emotional speech,”
Jianhua Tao, Yongguo Kang, and Aijun Li, · 2006
Earlier work this paper cites.
“Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory,”
Tomoki Toda, Alan W Black, and Keiichi Tokuda, · 2007
Earlier work this paper cites.
“Visualizing data using t-sne,”
Laurens van der Maaten and Geoffrey Hinton, · 2008
Earlier work this paper cites.
“Iemocap: Interactive emotional dyadic motion capture database,”
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan, · 2008
Earlier work this paper cites.
“Voice conversion using partial least squares regression,”
Elina Helander, Tuomas Virtanen, Jani Nurminen, and Moncef Gabbouj, · 2010
Earlier work this paper cites.
“Voice conversion using deep neural networks with layer-wise generative training,”
Ling-Hui Chen, Zhen-Hua Ling, Li-Juan Liu, and Li-Rong Dai, · 2014
Earlier work this paper cites.
“High-order sequence modeling using speaker-dependent recurrent temporal restricted boltzmann machines for voice conversion,”
Toru Nakashika, Tetsuya Takiguchi, and Yasuo Ariki, · 2014
Earlier work this paper cites.
“Exemplar-based emotional voice conversion using non-negative matrix factorization,”
Ryo Aihara, Reina Ueda, Tetsuya Takiguchi, and Yasuo Ariki, · 2014
Earlier work this paper cites.
“Emotional voice conversion using deep neural networks with mcc and f0 features,”
Zhaojie Luo, Tetsuya Takiguchi, and Yasuo Ariki, · 2016
Cited alongside, same era.
“Deep bidirectional lstm modeling of timbre and prosody for emotional voice conversion,”
Huaiping Ming, Dongyan Huang, Lei Xie, Jie Wu, Minghui Dong, and Haizhou Li, · 2016
Cited alongside, same era.
“World: a vocoder-based high-quality speech synthesis system for real-time applications,”
Masanori Morise, Fumiya Yokomori, and Kenji Ozawa, · 2016
Cited alongside, same era.
Chin-Cheng Hsu, Hsin-Te Hwang, Yi-Chiao Wu, Yu Tsao, and Hsin-Min Wang, · 2017
Cited alongside, same era.
“Adapting and controlling dnn-based speech synthesis using input codes,”
Hieu-Thi Luong, Shinji Takaki, Gustav Eje Henter, and Junichi Yamagishi, · 2017
Cited alongside, same era.
“Nonparallel emotional speech conversion,”
Jian Gao, Deep Chakraborty, Hamidou Tembine, and Olaitan Olaleye, · 2019
Later among the works it cites.
“Dnn-based emotion recognition based on bottleneck acoustic features and lexical features,”
Eesung Kim and Jong Won Shin, · 2019
Later among the works it cites.
“Teacher-student training for robust tacotron-based tts,”
Rui Liu, Berrak Sisman, Jingdong Li, Feilong Bao, Guanglai Gao, and Haizhou Li, · 2020
Closest in time.
“An overview of voice conversion and its challenges: From statistical modeling to deep learning,”
Berrak Sisman, Junichi Yamagishi, Simon King, and Haizhou Li, · 2020
Closest in time.
“Transforming Spectrum and Prosody for Emotional Voice Conversion with Non-Parallel Training Data,”
Kun Zhou, Berrak Sisman, and Haizhou Li, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Voice conversion for emotional speech: Rule-based synthesis with degree of emotion controllable in dimensional space,”
Yawen Xue, Yasuhiro Hamada, and Masato Akagi, · 2018
Cited alongside, same era.
“Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,”
Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ-Skerry Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Ye Jia, Fei Ren, and Rif A Saurous, · 2018
Cited alongside, same era.
“Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,”
RJ Skerry-Ryan, Eric Battenberg, Ying Xiao, Yuxuan Wang, Daisy Stanton, Joel Shor, Ron Weiss, Rob Clark, and Rif A Saurous, · 2018
Cited alongside, same era.
“3-d convolutional recurrent neural networks with attention model for speech emotion recognition,”
Mingyi Chen, Xuanji He, Jing Yang, and Han Zhang, · 2018
Cited alongside, same era.
“Group sparse representation with wavenet vocoder adaptation for spectrum and prosody conversion,”
Berrak Sisman, Mingyang Zhang, and Haizhou Li, · 2019
Cited alongside, same era.
“On the study of generative adversarial networks for cross-lingual voice conversion,”
Berrak Sisman, Mingyang Zhang, Minghui Dong, and Haizhou Li, · 2019
Cited alongside, same era.
“Sequence-to-sequence modelling of f0 for speech emotion conversion,”
Carl Robinson, Nicolas Obin, and Axel Roebel, · 2019
Cited alongside, same era.
Ravi Shankar, Jacob Sager, and Archana Venkataraman, · 2020
Closest in time.
“Converting Anyone’s Emotion: Towards Speaker-Independent Emotional Voice Conversion,”
Kun Zhou, Berrak Sisman, Mingyang Zhang, and Haizhou Li, · 2020
Closest in time.
“A review on five recent and near-future developments in computational processing of emotion in the human voice,”
Dagmar M Schuller and Björn W Schuller, · 2020
Closest in time.
“Deep representation learning in speech processing: Challenges, recent advances, and future trends,”
Siddique Latif, Rajib Rana, Sara Khalifa, Raja Jurdak, Junaid Qadir, and Björn W Schuller, · 2020
Closest in time.
“Interactive text-to-speech via semi-supervised style transfer learning,”
Yang Gao, Weiyi Zheng, Zhaojun Yang, Thilo Kohler, Christian Fuegen, and Qing He, · 2020
Closest in time.
“Emotional speech synthesis with rich and granularized control,”
Se-Yun Um, Sangshin Oh, Kyungguen Byun, Inseon Jang, ChungHyun Ahn, and Hong-Goo Kang, · 2020
Closest in time.
“Expressive tts training with frame and style reconstruction loss,”
Rui Liu, Berrak Sisman, Guanglai Gao, and Haizhou Li, · 2020
Closest in time.