Fetching the paper…
Reading the bibliography…
Emotional voice conversion (EVC) seeks to convert the emotional state of an utterance while preserving the linguistic content and speaker identity.
J. R. Davitz, M. Beldoch, S. Blau, L. Dimitrovsky, E. Levitt, P. Kempner Levy et al. , “Personality, perceptual, and cognitive correlates of emotional sensitivity,” The communication of emotional meaning , pp. 57–68, 1964
1964
Earlier work this paper cites.
J. A. Russell, “A circumplex model of affect.” Journal of personality and social psychology , vol. 39, no. 6, p. 1161, 1980
1980
Earlier work this paper cites.
W. F. Johnson, R. N. Emde, K. R. Scherer, and M. D. Klinnert, “Recognition of emotion from vocal cues,” Archives of General Psychiatry , vol. 43, no. 3, pp. 280–283, 1986
1986
Earlier work this paper cites.
C. M. Whissell, “The dictionary of affect in language,” in The measurement of emotions . Elsevier, 1989, pp. 113–131
1989
Earlier work this paper cites.
W. S. Sarle, “Algorithms for clustering data,” 1990
1990
Earlier work this paper cites.
N. H. Frijda, A. Ortony, J. Sonnemans, and G. L. Clore, “The complexity of intensity: Issues concerning the structure of emotion intensity.” 1992
1992
Earlier work this paper cites.
P. Ekman, “An argument for basic emotions,” Cognition & emotion , 1992
1992
Earlier work this paper cites.
D. W. Hosmer and S. Lemeshow, “Confidence interval estimation of interaction,” Epidemiology , pp. 452–456, 1992
1992
Earlier work this paper cites.
J. R. Averill and T. A. More, “Happiness.” 1993
1993
Earlier work this paper cites.
I. R. Murray and J. L. Arnott, “Toward the simulation of emotion in synthetic speech: A review of the literature on human vocal emotion,” The Journal of the Acoustical Society of America , vol. 93, no. 2, pp. 1097–1108, 1993
1993
Earlier work this paper cites.
R. Kubichek, “Mel-cepstral distance measure for objective speech quality assessment,” in Proceedings of IEEE Pacific Rim Conference on Communications Computers and Signal Processing , vol. 1. IEEE, 1993, pp. 125–128
1993
Earlier work this paper cites.
A. Kain and M. W. Macon, “Spectral voice conversion for text-to-speech synthesis,” in Proceedings of the 1998 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP’98 (Cat. No. 98CH36181) , vol. 1. IEEE, 1998, pp. 285–288
1998
Earlier work this paper cites.
R. A. Thisted, “What is a p-value,” Departments of Statistics and Health Studies , 1998
1998
Earlier work this paper cites.
J. W. Brehm, “The intensity of emotion,” Personality and social psychology review , vol. 3, no. 1, pp. 2–22, 1999
1999
Earlier work this paper cites.
P. N. Juslin and P. Laukka, “Impact of intended emotion intensity on cue utilization and decoding accuracy in vocal expression of emotion.” Emotion , vol. 1, no. 4, p. 381, 2001
2001
Earlier work this paper cites.
A. Black, P. Taylor, R. Caley, R. Clark, K. Richmond, S. King, V. Strom, and H. Zen, “The festival speech synthesis system, version 1.4.2,” Unpublished document available via http://www.cstr.ed.ac.uk/projects/festival.html , 2001
2001
Earlier work this paper cites.
J. Hirschberg, “Communication and prosody: Functional aspects of prosody,” Speech Communication , vol. 36, no. 1-2, pp. 31–43, 2002
2002
Earlier work this paper cites.
U. Maulik and S. Bandyopadhyay, “Performance evaluation of some clustering algorithms and validity indices,” IEEE Transactions on pattern analysis and machine intelligence , vol. 24, no. 12, pp. 1650–1654, 2002
2002
Earlier work this paper cites.
Y. Zhao and G. Karypis, “Evaluation of hierarchical clustering algorithms for document datasets,” in Proceedings of the eleventh international conference on Information and knowledge management , 2002, pp. 515–524
2002
Earlier work this paper cites.
H. Kawanami, Y. Iwami, T. Toda, H. Saruwatari, and K. Shikano, “Gmm-based voice conversion applied to emotional speech synthesis,” 2003
2003
Earlier work this paper cites.
O. Pierre-Yves, “The production and recognition of emotions in speech: features and algorithms,” International Journal of Human-Computer Studies , vol. 59, no. 1-2, pp. 157–183, 2003
2003
Earlier work this paper cites.
K. R. Scherer, “Vocal communication of emotion: A review of research paradigms,” Speech communication , vol. 40, no. 1-2, pp. 227–256, 2003
2003
Earlier work this paper cites.
J. Kominek and A. W. Black, “The cmu arctic speech databases,” in Fifth ISCA workshop on speech synthesis , 2004
2004
Earlier work this paper cites.
D. Erickson, “Expressive speech: Production, perception and application to speech synthesis,” Acoustical science and technology , vol. 26, no. 4, pp. 317–325, 2005
2005
Earlier work this paper cites.
J. Tao, Y. Kang, and A. Li, “Prosody conversion from neutral speech to emotional speech,” IEEE Transactions on Audio, Speech, and Language Processing , 2006
2006
Earlier work this paper cites.
D. Ververidis and C. Kotropoulos, “Emotional speech recognition: Resources, features, and methods,” Speech communication , vol. 48, no. 9, pp. 1162–1181, 2006
2006
Earlier work this paper cites.
V. Ferrari and A. Zisserman, “Learning visual attributes,” Advances in neural information processing systems , vol. 20, pp. 433–440, 2007
2007
Earlier work this paper cites.
O. Chapelle, “Training a support vector machine in the primal,” Neural computation , vol. 19, no. 5, pp. 1155–1178, 2007
2007
Earlier work this paper cites.
A. Rosenberg and J. Hirschberg, “V-measure: A conditional entropy-based external cluster evaluation measure,” in Proceedings of the 2007 joint conference on empirical methods in natural language processing and computational natural language learning (EMNLP-CoNLL) , 2007, pp. 410–420
2007
Earlier work this paper cites.
M. J. Owren and J.-A. Bachorowski, “Measuring emotion-related vocal acoustics,” Handbook of emotion elicitation and assessment , pp. 239–266, 2007
2007
Earlier work this paper cites.
M. Müller, “Dynamic time warping,” Information retrieval for music and motion , pp. 69–84, 2007
2007
Earlier work this paper cites.
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “Iemocap: Interactive emotional dyadic motion capture database,” Language resources and evaluation , vol. 42, no. 4, p. 335, 2008
2008
Earlier work this paper cites.
L. v. d. Maaten and G. Hinton, “Visualizing data using t-sne,” Journal of machine learning research , vol. 9, no. Nov, pp. 2579–2605, 2008
2008
Earlier work this paper cites.
J. Wang, R. Dou, Z. Yan, Z. Wang, and Z. Zhang, “Exploration of analyzing emotion strength in speech signal,” in 2009 Chinese Conference on Pattern Recognition . IEEE, 2009, pp. 1–4
2009
Earlier work this paper cites.
A. Rosenberg and J. Hirschberg, “Detecting pitch accents at the word, syllable and vowel level,” in Proceedings of Human Language Technologies: The 2009 Annual Conference of the North American Chapter of the Association for Computational Linguistics, Companion Volume: Short Papers , 2009, pp. 81–84
2009
Earlier work this paper cites.
Z. Inanoglu and S. Young, “Data-driven emotion conversion in spoken english,” Speech Communication , 2009
2009
Earlier work this paper cites.
B. Schuller, S. Steidl, and A. Batliner, “The interspeech 2009 emotion challenge,” in Tenth Annual Conference of the International Speech Communication Association , 2009
2009
Earlier work this paper cites.
J. Pittermann, A. Pittermann, and W. Minker, Handling emotions in human-computer dialogues . Springer, 2010
2010
Earlier work this paper cites.
D. Wu, T. D. Parsons, and S. S. Narayanan, “Acoustic feature analysis in speech emotion primitives estimation,” in Eleventh Annual Conference of the International Speech Communication Association , 2010
2010
Earlier work this paper cites.
B. McFee and G. R. Lanckriet, “Metric learning to rank,” in ICML , 2010
2010
Earlier work this paper cites.
Y. Liu, Z. Li, H. Xiong, X. Gao, and J. Wu, “Understanding of internal clustering validation measures,” in 2010 IEEE international conference on data mining . IEEE, 2010, pp. 911–916
2010
Earlier work this paper cites.
F. Eyben, M. Wöllmer, and B. Schuller, “Opensmile: the munich versatile and fast open-source audio feature extractor,” in Proceedings of the 18th ACM international conference on Multimedia , 2010, pp. 1459–1462
2010
Earlier work this paper cites.
R. Levitan and J. Hirschberg, “Measuring acoustic-prosodic entrainment with respect to multiple levels and dimensions,” in Twelfth Annual Conference of the International Speech Communication Association , 2011
2011
Earlier work this paper cites.
Y. Xu, “Speech prosody: A methodological review,” Journal of Speech Sciences , vol. 1, no. 1, pp. 85–115, 2011
2011
Earlier work this paper cites.
D. Parikh and K. Grauman, “Relative attributes,” in 2011 International Conference on Computer Vision . IEEE, 2011, pp. 503–510
2011
Earlier work this paper cites.
S. Ramakrishnan, Speech Enhancement, Modeling and Recognition-Algorithms and Applications . BoD–Books on Demand, 2012
2012
Earlier work this paper cites.
R. Aihara, R. Takashima, T. Takiguchi, and Y. Ariki, “Gmm-based emotional voice conversion using spectrum and prosody features,” American Journal of Signal Processing , 2012
2012
Cited alongside, same era.
A. Kovashka, D. Parikh, and K. Grauman, “Whittlesearch: Image search with relative attribute feedback,” in 2012 IEEE Conference on Computer Vision and Pattern Recognition . IEEE, 2012, pp. 2973–2980
2012
Cited alongside, same era.
B. Schuller and A. Batliner, Computational paralinguistics: emotion, affect and personality in speech and language processing . John Wiley & Sons, 2013
2013
Cited alongside, same era.
2013
Cited alongside, same era.
Y.-J. Zhang, S. Pan, L. He, and Z.-H. Ling, “Learning latent representations for style control and transfer in end-to-end speech synthesis,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6945–6949
2019
Later among the works it cites.
T. Kenter, V. Wan, C.-A. Chan, R. Clark, and J. Vit, “Chive: Varying prosody in speech synthesis with a linguistically driven dynamic hierarchical conditional variational network,” in International Conference on Machine Learning . PMLR, 2019, pp. 3331–3340
2019
Later among the works it cites.
X. Zhu, S. Yang, G. Yang, and L. Xie, “Controlling emotion strength with relative attribute for end-to-end speech synthesis,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 192–199
2019
Later among the works it cites.
D. Ong, Z. Wu, Z.-X. Tan, M. Reddan, I. Kahhale, A. Mattek, and J. Zaki, “Modeling emotion in complex stories: the stanford emotional narratives dataset,” IEEE Transactions on Affective Computing , 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2013
Cited alongside, same era.
R. Aihara, R. Ueda, T. Takiguchi, and Y. Ariki, “Exemplar-based emotional voice conversion using non-negative matrix factorization,” in APSIPA ASC . IEEE, 2014
2014
Cited alongside, same era.
D. T. Larose and C. D. Larose, Discovering knowledge in data: an introduction to data mining . John Wiley & Sons, 2014, vol. 4
2014
Cited alongside, same era.
D. Bahdanau, K. H. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in 3rd International Conference on Learning Representations, ICLR 2015 , 2015
2015
Cited alongside, same era.
Y. Fan, Y. Qian, F. K. Soong, and L. He, “Multi-speaker modeling and speaker adaptation for dnn-based tts synthesis,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2015, pp. 4475–4479
2015
Cited alongside, same era.
G. Koch, R. Zemel, R. Salakhutdinov et al. , “Siamese neural networks for one-shot image recognition,” in ICML deep learning workshop , vol. 2. Lille, 2015, p. 0
2015
Cited alongside, same era.
Z. Zhang, C. Wang, B. Xiao, W. Zhou, and S. Liu, “Robust relative attributes for human action recognition,” Pattern Analysis and Applications , vol. 18, no. 1, pp. 157–171, 2015
2015
Cited alongside, same era.
H. Muthusamy, K. Polat, and S. Yaacob, “Improved emotion recognition using gaussian mixture model and extreme learning machine in speech and glottal signals,” Mathematical Problems in Engineering , vol. 2015, 2015
2015
Cited alongside, same era.
2019
Later among the works it cites.
E. Kim and J. W. Shin, “Dnn-based emotion recognition based on bottleneck acoustic features and lexical features,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6720–6724
2019
Later among the works it cites.
O. Kwon, I. Jang, C. Ahn, and H.-G. Kang, “An effective style token weight control technique for end-to-end emotional speech synthesis,” IEEE Signal Processing Letters , vol. 26, no. 9, pp. 1383–1387, 2019
2019
Later among the works it cites.
B. Sisman, J. Yamagishi, S. King, and H. Li, “An overview of voice conversion and its challenges: From statistical modeling to deep learning,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2020
2020
Later among the works it cites.
K. Matsumoto, S. Hara, and M. Abe, “Controlling the Strength of Emotions in Speech-Like Emotional Sound Generated by WaveNet,” in Proc. Interspeech 2020 , 2020, pp. 3421–3425
2020
Later among the works it cites.
S.-Y. Um, S. Oh, K. Byun, I. Jang, C. Ahn, and H.-G. Kang, “Emotional speech synthesis with rich and granularized control,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7254–7258
2020
Later among the works it cites.
K. Zhou, B. Sisman, and H. Li, “Transforming Spectrum and Prosody for Emotional Voice Conversion with Non-Parallel Training Data,” in Proc. Odyssey 2020 The Speaker and Language Recognition Workshop , 2020, pp. 230–237
2020
Later among the works it cites.
2020
Later among the works it cites.
G. Rizos, A. Baird, M. Elliott, and B. Schuller, “Stargan for emotional speech conversion: Validated by data augmentation of end-to-end emotion recognition,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 3502–3506
2020
Later among the works it cites.
K. Zhou, B. Sisman, M. Zhang, and H. Li, “Converting Anyone’s Emotion: Towards Speaker-Independent Emotional Voice Conversion,” in Proc. Interspeech 2020 , 2020, pp. 3416–3420
2020
Later among the works it cites.
R. Shankar, H.-W. Hsieh, N. Charon, and A. Venkataraman, “Multi-speaker emotion conversion via latent variable regularization and a chained encoder-decoder-predictor network,” Proc. Interspeech 2020 , pp. 3391–3395, 2020
2020
Later among the works it cites.
K. Qian, Z. Jin, M. Hasegawa-Johnson, and G. J. Mysore, “F0-consistent many-to-many non-parallel voice conversion via conditional autoencoder,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6284–6288
2020
Later among the works it cites.
H. Kameoka, W.-C. Huang, K. Tanaka, T. Kaneko, N. Hojo, and T. Toda, “Many-to-many voice transformer network,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 656–670, 2020
2020
Later among the works it cites.
T.-H. Kim, S. Cho, S. Choi, S. Park, and S.-Y. Lee, “Emotional voice conversion using multitask learning with text-to-speech,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7774–7778
2020
Later among the works it cites.
H. Kameoka, K. Tanaka, D. Kwaśny, T. Kaneko, and N. Hojo, “Convs2s-vc: Fully convolutional sequence-to-sequence voice conversion,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 1849–1863, 2020
2020
Later among the works it cites.
D. M. Schuller and B. W. Schuller, “A review on five recent and near-future developments in computational processing of emotion in the human voice,” Emotion Review , p. 1754073919898526, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
J. Y. Lee, S. J. Cheon, B. J. Choi, and N. S. Kim, “Memory attention: Robust alignment using gating mechanism for end-to-end speech synthesis,” IEEE Signal Processing Letters , vol. 27, pp. 2004–2008, 2020
2020
Later among the works it cites.
R. Yamamoto, E. Song, and J.-M. Kim, “Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6199–6203
2020
Later among the works it cites.
J.-X. Zhang, Z.-H. Ling, and L.-R. Dai, “Recognition-Synthesis Based Non-Parallel Voice Conversion with Adversarial Learning,” in Proc. Interspeech 2020 , 2020, pp. 771–775
2020
Later among the works it cites.
U. Tiwari, M. Soni, R. Chakraborty, A. Panda, and S. K. Kopparapu, “Multi-conditioning and data augmentation using generative noise model for speech emotion recognition in noisy conditions,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7194–7198
2020
Later among the works it cites.
A. Rosenberg and J. Hirschberg, “Prosodic aspects of the attractive voice,” in Voice Attractiveness . Springer, 2021, pp. 17–40
2021
Later among the works it cites.
H. Choi and M. Hahn, “Sequence-to-sequence emotional voice conversion with strength control,” IEEE Access , vol. 9, pp. 42 674–42 687, 2021
2021
Later among the works it cites.
K. Zhou, B. Sisman, and H. Li, “Vaw-gan for disentanglement and recomposition of emotional elements in speech,” in 2021 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2021, pp. 415–422
2021
Later among the works it cites.
2021
Later among the works it cites.
K. Zhou, B. Sisman, R. Liu, and H. Li, “Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 920–924
2021
Later among the works it cites.
R. Liu, B. Sisman, G. lai Gao, and H. Li, “Expressive tts training with frame and style reconstruction loss,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2021
2021
Later among the works it cites.
T. Hayashi, W.-C. Huang, K. Kobayashi, and T. Toda, “Non-autoregressive sequence-to-sequence voice conversion,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 7068–7072
2021
Later among the works it cites.
2021
Later among the works it cites.
W.-C. Huang, T. Hayashi, Y.-C. Wu, H. Kameoka, and T. Toda, “Pretraining techniques for sequence-to-sequence voice conversion,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 745–755, 2021
2021
Later among the works it cites.
S. Liu, Y. Cao, D. Wang, X. Wu, X. Liu, and H. Meng, “Any-to-many voice conversion with location-relative sequence-to-sequence modeling,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 1717–1728, 2021
2021
Later among the works it cites.
K. Zhou, B. Sisman, and H. Li, “Limited Data Emotional Voice Conversion Leveraging Text-to-Speech: Two-Stage Sequence-to-Sequence Training,” in Proc. Interspeech 2021 , 2021, pp. 811–815
2021
Later among the works it cites.
X. Li, C. Song, J. Li, Z. Wu, J. Jia, and H. Meng, “Towards Multi-Scale Style Control for Expressive Speech Synthesis,” in Proc. Interspeech 2021 , 2021, pp. 4673–4677
2021
Later among the works it cites.
D. Tan and T. Lee, “Fine-Grained Style Modeling, Transfer and Prediction in Text-to-Speech Synthesis via Phone-Level Content-Style Disentanglement,” in Proc. Interspeech 2021 , 2021, pp. 4683–4687
2021
Later among the works it cites.
Y. Lei, S. Yang, and L. Xie, “Fine-grained emotion strength transfer, control and prediction for emotional speech synthesis,” in 2021 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2021, pp. 423–430
2021
Later among the works it cites.
S. Latif, R. Rana, S. Khalifa, R. Jurdak, J. Qadir, and B. W. Schuller, “Survey of deep representation learning for speech emotion recognition,” IEEE Transactions on Affective Computing , 2021
2021
Later among the works it cites.
R. Liu, B. Sisman, and H. Li, “Reinforcement Learning for Emotional Text-to-Speech Synthesis with Improved Emotion Discriminability,” in Proc. Interspeech 2021 , 2021, pp. 4648–4652
2021
Later among the works it cites.
T. Li, S. Yang, L. Xue, and L. Xie, “Controllable emotion transfer for end-to-end speech synthesis,” in 2021 12th International Symposium on Chinese Spoken Language Processing (ISCSLP) . IEEE, 2021, pp. 1–5
2021
Later among the works it cites.
X. Cai, D. Dai, Z. Wu, X. Li, J. Li, and H. Meng, “Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 5734–5738
2021
Later among the works it cites.
B. Schnell, G. Huybrechts, B. Perz, T. Drugman, and J. Lorenzo-Trueba, “Emocat: Language-agnostic emotional voice conversion,” 2021
2021
Later among the works it cites.
B. J. Abbaschian, D. Sierra-Sosa, and A. Elmaghraby, “Deep learning techniques for speech emotion recognition, from databases to models,” Sensors , vol. 21, no. 4, p. 1249, 2021
2021
Later among the works it cites.
K. Zhou, B. Sisman, R. Liu, and H. Li, “Emotional voice conversion: Theory, databases and esd,” Speech Communication , vol. 137, pp. 1–18, 2022
2022
Closest in time.