Fetching the paper…
Reading the bibliography…
Emotional voice conversion aims to convert the spectrum and prosody to change the emotional patterns of speech, while preserving the speaker identity and linguistic content.
“A voice conversion framework with tandem feature sparse representation and speaker-adapted wavenet vocoder.,”
Berrak Sisman, Mingyang Zhang, and Haizhou Li, · 1982
Earlier work this paper cites.
“Voice conversion through vector quantization,”
Masanobu Abe, Satoshi Nakamura, Kiyohiro Shikano, and Hisao Kuwabara, · 1990
Earlier work this paper cites.
“Vocal cues in emotion encoding and decoding,”
Klaus R Scherer, Rainer Banse, Harald G Wallbott, and Thomas Goldbeck, · 1991
Earlier work this paper cites.
“Speaker adaptation and voice conversion by codebook mapping,”
Kiyohiro Shikano, Satoshi Nakamura, and Masanobu Abe, · 1991
Earlier work this paper cites.
“Algorithms for non-negative matrix factorization,”
Daniel D Lee and H Sebastian Seung, · 2001
Earlier work this paper cites.
“Emotional speech synthesis: A review,”
Marc Schröder, · 2001
Earlier work this paper cites.
“A corpus-based speech synthesis system with emotion,”
Akemi Iida, Nick Campbell, Fumito Higuchi, and Michiaki Yasumura, · 2003
Earlier work this paper cites.
“Estimation of the parameters of the quantitative intonation model with continuous wavelet analysis,”
Hans Kruschke and Michael Lenz, · 2003
Earlier work this paper cites.
“A fast learning algorithm for deep belief nets,”
Geoffrey E Hinton, Simon Osindero, and Yee-Whye Teh, · 2006
Earlier work this paper cites.
“Prosody conversion from neutral speech to emotional speech,”
Jianhua Tao, Yongguo Kang, and Aijun Li, · 2006
Earlier work this paper cites.
“Decomposition of pitch curves in the general superpositional intonation model,”
Taniya Mishra, Jan Van Santen, and Esther Klabbers, · 2006
Earlier work this paper cites.
“Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory,”
Tomoki Toda, Alan W Black, and Keiichi Tokuda, · 2007
Earlier work this paper cites.
“Multilevel parametric-base f0 model for speech synthesis,”
Javier Latorre and Masami Akamine, · 2008
Earlier work this paper cites.
“Hierarchical prosody conversion using regression-based clustering for emotional speech synthesis,”
Chung-Hsien Wu, Chi-Chun Hsia, Chung-Han Lee, and Mai-Chun Lin, · 2009
Earlier work this paper cites.
“Data-driven emotion conversion in spoken english,”
Zeynep Inanoglu and Steve Young, · 2009
Earlier work this paper cites.
“Voice conversion using partial least squares regression,”
Elina Helander, Tuomas Virtanen, Jani Nurminen, and Moncef Gabbouj, · 2010
Earlier work this paper cites.
“Spectral mapping using artificial neural networks for voice conversion,”
Srinivas Desai, Alan W Black, B Yegnanarayana, and Kishore Prahallad, · 2010
Earlier work this paper cites.
“Speech prosody: A methodological review,”
Yi Xu, · 2011
Earlier work this paper cites.
Speech Enhancement, Modeling and Recognition-Algorithms and Applications
S Ramakrishnan, · 2012
Earlier work this paper cites.
“Gmm-based emotional voice conversion using spectrum and prosody features,”
Ryo Aihara, Ryoichi Takashima, Tetsuya Takiguchi, and Yasuo Ariki, · 2012
Earlier work this paper cites.
“Wavelets for intonation modeling in hmm speech synthesis,”
Antti Santeri Suni, Daniel Aalto, Tuomo Raitio, Paavo Alku, Martti Vainio, et al., · 2013
Earlier work this paper cites.
“Continuous wavelet transform for analysis of speech prosody,”
Martti Vainio, Antti Suni, Daniel Aalto, et al., · 2013
Earlier work this paper cites.
“Exemplar-based emotional voice conversion using non-negative matrix factorization,”
Ryo Aihara, Reina Ueda, Tetsuya Takiguchi, and Yasuo Ariki, · 2014
Cited alongside, same era.
“Exemplar-based sparse representation with residual compensation for voice conversion,”
Zhizheng Wu, Tuomas Virtanen, Eng Siong Chng, and Haizhou Li, · 2014
Cited alongside, same era.
“Voice conversion using deep neural networks with layer-wise generative training,”
Ling-Hui Chen, Zhen-Hua Ling, Li-Juan Liu, and Li-Rong Dai, · 2014
Cited alongside, same era.
“High-order sequence modeling using speaker-dependent recurrent temporal restricted boltzmann machines for voice conversion,”
Toru Nakashika, Tetsuya Takiguchi, and Yasuo Ariki, · 2014
Cited alongside, same era.
“Hierarchical modeling of f0 contours for voice conversion,”
Gerard Sanchez, Hanna Silen, Jani Nurminen, and Moncef Gabbouj, · 2014
Cited alongside, same era.
“Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks,”
Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, and Dimitris N Metaxas, · 2017
Later among the works it cites.
“Unpaired image-to-image translation using cycle-consistent adversarial networks,”
Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros, · 2017
Later among the works it cites.
“Conditional cyclegan for attribute guided face image generation,”
Yongyi Lu, Yu-Wing Tai, and Chi-Keung Tang, · 2017
Later among the works it cites.
“Nonparallel emotional speech conversion,”
Jian Gao, Deep Chakraborty, Hamidou Tembine, and Olaitan Olaleye, · 2018
Later among the works it cites.
“Cyclegan-vc: Non-parallel voice conversion using cycle-consistent adversarial networks,”
Takuhiro Kaneko and Hirokazu Kameoka, · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio, · 2014
Cited alongside, same era.
“Emotional facial expression transfer based on temporal restricted boltzmann machines,”
Shuojun Liu, Dong-Yan Huang, Weisi Lin, Minghui Dong, Haizhou Li, and Ee Ping Ong, · 2014
Cited alongside, same era.
“Fundamental frequency modeling using wavelets for emotional voice conversion,”
Huaiping Ming, Dongyan Huang, Minghui Dong, Haizhou Li, Lei Xie, and Shaofei Zhang, · 2015
Cited alongside, same era.
“On the use of i-vectors and average voice model for voice conversion without parallel data,”
Jie Wu, Zhizheng Wu, and Lei Xie, · 2016
Cited alongside, same era.
“Voice conversion from non-parallel corpora using variational auto-encoder,”
Chin-Cheng Hsu, Hsin-Te Hwang, Yi-Chiao Wu, Yu Tsao, and Hsin-Min Wang, · 2016
Cited alongside, same era.
“Phonetic posteriorgrams for many-to-one voice conversion without parallel data training,”
Lifa Sun, Kun Li, Hao Wang, Shiyin Kang, and Helen Meng, · 2016
Cited alongside, same era.
“Emotional voice conversion using deep neural networks with mcc and f0 features,”
Zhaojie Luo, Tetsuya Takiguchi, and Yasuo Ariki, · 2016
Cited alongside, same era.
Later among the works it cites.
“Investigating different representations for modeling and controlling multiple emotions in dnn-based speech synthesis,”
Jaime Lorenzo-Trueba, Gustav Eje Henter, Shinji Takaki, Junichi Yamagishi, Yosuke Morino, and Yuta Ochiai, · 2018
Later among the works it cites.
“Voice conversion for emotional speech: Rule-based synthesis with degree of emotion controllable in dimensional space,”
Yawen Xue, Yasuhiro Hamada, and Masato Akagi, · 2018
Later among the works it cites.
“Wavelet analysis of speaker dependent and independent prosody for voice conversion.,”
Berrak Sisman and Haizhou Li, · 2018
Later among the works it cites.
“Adaptive wavenet vocoder for residual compensation in gan-based voice conversion,”
Berrak Sisman, Mingyang Zhang, Sakti Sakriani, Haizhou Li, and Satoshi Nakamura, · 2018
Later among the works it cites.
“Phonetically aware exemplar-based prosody transformation,”
Berrak Sisman, Grandee Lee, and Haizhou Li, · 2018
Later among the works it cites.
“On the study of generative adversarial networks for cross-lingual voice conversion,”
Berrak Sisman, Mingyang Zhang, Minghui Dong, and Haizhou Li, · 2019
Later among the works it cites.
“Sequence-to-sequence modelling of f0 for speech emotion conversion,”
Carl Robinson, Nicolas Obin, and Axel Roebel, · 2019
Later among the works it cites.
“Group sparse representation with wavenet vocoder adaptation for spectrum and prosody conversion,”
Berrak Sisman, Mingyang Zhang, and Haizhou Li, · 2019
Later among the works it cites.
“Emotional voice conversion using dual supervised adversarial networks with continuous wavelet transform f0 features,”
Zhaojie Luo, Jinhui Chen, Tetsuya Takiguchi, and Yasuo Ariki, · 2019
Later among the works it cites.
“Semantically consistent hierarchical text to fashion image synthesis with an enhanced-attentional generative adversarial network,”
Kenan Emir Ak, Joo Hwee Lim, Jo Yew Tham, and Ashraf Kassim, · 2019
Later among the works it cites.
“Attribute manipulation generative adversarial networks for fashion images,”
Kenan E Ak, Joo Hwee Lim, Jo Yew Tham, and Ashraf A Kassim, · 2019
Later among the works it cites.
“SINGAN: Singing voice conversion with generative adversarial networks,”
Berrak Sisman, Karthika Vijayan, Minghui Dong, and Haizhou Li, · 2019
Later among the works it cites.
“Cyclegan-vc2: Improved cyclegan-based non-parallel voice conversion,”
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, and Nobukatsu Hojo, · 2019
Later among the works it cites.
“Wavetts: Tacotron-based tts with joint time-frequency domain loss,”
Rui Liu, Berrak Sisman, Feilong Bao, Guanglai Gao, and Haizhou Li, · 2020
Closest in time.
“Teacher-student training for robust tacotron-based tts,”
Rui Liu, Berrak Sisman, Jingdong Li, Feilong Bao, Guanglai Gao, and Haizhou Li, · 2020
Closest in time.
“Semantically consistent text to fashion image synthesis with an enhanced attentional generative adversarial network,”
Kenan E Ak, Joo Hwee Lim, Jo Yew Tham, and Ashraf A Kassim, · 2020
Closest in time.