Fetching the paper…
Reading the bibliography…
Emotional speech synthesis aims to synthesize human voices with various emotional effects.
G. A. Miller, “The magical number seven, plus or minus two: Some limits on our capacity for processing information.” Psychological review , vol. 63, no. 2, p. 81, 1956
1956
Earlier work this paper cites.
J. A. Russell, “A circumplex model of affect.” Journal of personality and social psychology , vol. 39, no. 6, p. 1161, 1980
1980
Earlier work this paper cites.
W. F. Johnson, R. N. Emde, K. R. Scherer, and M. D. Klinnert, “Recognition of emotion from vocal cues,” Archives of General Psychiatry , vol. 43, no. 3, pp. 280–283, 1986
1986
Earlier work this paper cites.
C. M. Whissell, “The dictionary of affect in language,” in The measurement of emotions . Elsevier, 1989, pp. 113–131
1989
Earlier work this paper cites.
R. Plutchik, The emotions . University Press of America, 1991
1991
Earlier work this paper cites.
N. H. Frijda, A. Ortony, J. Sonnemans, and G. L. Clore, “The complexity of intensity: Issues concerning the structure of emotion intensity.” 1992
1992
Earlier work this paper cites.
P. Ekman, “An argument for basic emotions,” Cognition & emotion , 1992
1992
Earlier work this paper cites.
R. Kubichek, “Mel-cepstral distance measure for objective speech quality assessment,” in Proceedings of IEEE Pacific Rim Conference on Communications Computers and Signal Processing , vol. 1. IEEE, 1993, pp. 125–128
1993
Earlier work this paper cites.
P. E. Ekman and R. J. Davidson, The nature of emotion: Fundamental questions. Oxford University Press, 1994
1994
Earlier work this paper cites.
L. F. Barrett, “Discrete emotions or dimensions? the role of valence focus and arousal focus,” Cognition & Emotion , vol. 12, no. 4, pp. 579–599, 1998
1998
Earlier work this paper cites.
M. Schröder, “Emotional speech synthesis: A review,” in Seventh European Conference on Speech Communication and Technology . Citeseer, 2001
2001
Earlier work this paper cites.
R. Plutchik, “The nature of emotions: Human emotions have deep evolutionary roots, a fact that may explain their complexity and provide tools for clinical practice,” American scientist , vol. 89, no. 4, pp. 344–350, 2001
2001
Earlier work this paper cites.
A. Black, P. Taylor, R. Caley, R. Clark, K. Richmond, S. King, V. Strom, and H. Zen, “The festival speech synthesis system, version 1.4.2,” Unpublished document available via http://www.cstr.ed.ac.uk/projects/festival.html , 2001
2001
Earlier work this paper cites.
K. Tokuda, H. Zen, and A. W. Black, “An hmm-based speech synthesis system applied to english,” in IEEE speech synthesis workshop . IEEE Santa Monica, 2002, pp. 227–230
2002
Earlier work this paper cites.
S. Mozziconacci, “Prosody and emotions,” in Speech Prosody 2002, International Conference , 2002
2002
Earlier work this paper cites.
P. Williams and J. L. Aaker, “Can mixed emotions peacefully coexist?” Journal of consumer research , vol. 28, no. 4, pp. 636–649, 2002
2002
Earlier work this paper cites.
J. Yamagishi, K. Onishi, T. Masuko, and T. Kobayashi, “Modeling of various speaking styles and emotions for hmm-based speech synthesis,” in Eighth European Conference on Speech Communication and Technology , 2003
2003
Earlier work this paper cites.
H. Kawanami, Y. Iwami, T. Toda, H. Saruwatari, and K. Shikano, “Gmm-based voice conversion applied to emotional speech synthesis,” 2003
2003
Earlier work this paper cites.
P. C. Hogan, The mind and its stories: Narrative universals and human emotion . Cambridge University Press, 2003
2003
Earlier work this paper cites.
J. Fürnkranz and E. Hüllermeier, “Pairwise preference learning and ranking,” in European conference on machine learning . Springer, 2003, pp. 145–156
2003
Earlier work this paper cites.
M. Schroder, “Expressing degree of activation in synthetic speech,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 14, no. 4, pp. 1128–1136, 2006
2006
Earlier work this paper cites.
O. Chapelle, “Training a support vector machine in the primal,” Neural computation , vol. 19, no. 5, pp. 1155–1178, 2007
2007
Earlier work this paper cites.
M. J. Owren and J.-A. Bachorowski, “Measuring emotion-related vocal acoustics,” Handbook of emotion elicitation and assessment , pp. 239–266, 2007
2007
Earlier work this paper cites.
J. Latorre and M. Akamine, “Multilevel parametric-base f0 model for speech synthesis,” in Ninth Annual Conference of the International Speech Communication Association , 2008
2008
Earlier work this paper cites.
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “Iemocap: Interactive emotional dyadic motion capture database,” Language resources and evaluation , vol. 42, no. 4, p. 335, 2008
2008
Earlier work this paper cites.
J. Yamagishi, T. Kobayashi, Y. Nakano, K. Ogata, and J. Isogai, “Analysis of speaker adaptation algorithms for hmm-based speech synthesis and a constrained smaplr adaptation algorithm,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 17, no. 1, pp. 66–83, 2009
2009
Earlier work this paper cites.
B. Schuller, S. Steidl, and A. Batliner, “The interspeech 2009 emotion challenge,” in Tenth Annual Conference of the International Speech Communication Association , 2009
2009
Earlier work this paper cites.
J. Pittermann, A. Pittermann, and W. Minker, Handling emotions in human-computer dialogues . Springer, 2010
2010
Earlier work this paper cites.
D. Wu, T. D. Parsons, and S. S. Narayanan, “Acoustic feature analysis in speech emotion primitives estimation,” in Eleventh Annual Conference of the International Speech Communication Association , 2010
2010
Earlier work this paper cites.
F. Eyben, M. Wöllmer, and B. Schuller, “Opensmile: the munich versatile and fast open-source audio feature extractor,” in Proceedings of the 18th ACM international conference on Multimedia , 2010, pp. 1459–1462
2010
Earlier work this paper cites.
D. Bitouk, R. Verma, and A. Nenkova, “Class-level spectral features for emotion recognition,” Speech communication , vol. 52, no. 7-8, pp. 613–625, 2010
2010
Earlier work this paper cites.
Y. Miyamoto, Y. Uchida, and P. C. Ellsworth, “Culture and mixed emotions: co-occurrence of positive and negative emotions in japan and the united states.” Emotion , vol. 10, no. 3, p. 404, 2010
2010
Earlier work this paper cites.
——, “Further evidence for mixed emotions.” Journal of personality and social psychology , vol. 100, no. 6, p. 1095, 2011
2011
Earlier work this paper cites.
Y. Xu, “Speech prosody: A methodological review,” Journal of Speech Sciences , vol. 1, no. 1, pp. 85–115, 2011
2011
Earlier work this paper cites.
D. Parikh and K. Grauman, “Relative attributes,” in 2011 International Conference on Computer Vision . IEEE, 2011, pp. 503–510
2011
Earlier work this paper cites.
F. Eyben, S. Buchholz, N. Braunschweiler, J. Latorre, V. Wan, M. J. Gales, and K. Knill, “Unsupervised clustering of emotion and voice styles for expressive tts,” in ICASSP 2012 - 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2012, pp. 4009–4012
2012
Earlier work this paper cites.
R. Aihara, R. Takashima, T. Takiguchi, and Y. Ariki, “Gmm-based emotional voice conversion using spectrum and prosody features,” American Journal of Signal Processing , 2012
2012
Earlier work this paper cites.
P. M. Niedenthal and M. Brauer, “Social functionality of human emotion,” Annual review of psychology , vol. 63, pp. 259–285, 2012
2012
Earlier work this paper cites.
A. Kovashka, D. Parikh, and K. Grauman, “Whittlesearch: Image search with relative attribute feedback,” in 2012 IEEE Conference on Computer Vision and Pattern Recognition . IEEE, 2012, pp. 2973–2980
2012
Earlier work this paper cites.
J. P. Van Santen, R. Sproat, J. Olive, and J. Hirschberg, Progress in speech synthesis . Springer Science & Business Media, 2013
2013
Earlier work this paper cites.
H. Ze, A. Senior, and M. Schuster, “Statistical parametric speech synthesis using deep neural networks,” in 2013 ieee international conference on acoustics, speech and signal processing . IEEE, 2013, pp. 7962–7966
2013
Earlier work this paper cites.
R. Plutchik and H. Kellerman, Theories of emotion . Academic Press, 2013, vol. 1
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
Q. Fan, P. Gabbur, and S. Pankanti, “Relative attributes for large-scale abandoned object detection,” in Proceedings of the IEEE International Conference on Computer Vision , 2013, pp. 2736–2743
2013
Earlier work this paper cites.
A. Braniecka, E. Trzebińska, A. Dowgiert, and A. Wytykowska, “Mixed emotions and coping: The benefits of secondary emotions,” PloS one , vol. 9, no. 8, p. e103940, 2014
2014
Earlier work this paper cites.
J. T. Larsen and A. P. McGraw, “The case for mixed emotions,” Social and Personality Psychology Compass , vol. 8, no. 6, pp. 263–274, 2014
2014
Earlier work this paper cites.
H. P. Martinez, G. N. Yannakakis, and J. Hallam, “Don’t classify ratings of affect; rank them!” IEEE transactions on affective computing , vol. 5, no. 3, pp. 314–326, 2014
2014
Cited alongside, same era.
Y. Ohtani, Y. Nasu, M. Morita, and M. Akamine, “Emotional transplant in statistical speech synthesis based on emotion additive model,” in Sixteenth Annual Conference of the International Speech Communication Association , 2015
2015
Cited alongside, same era.
R. Berrios, P. Totterdell, and S. Kellett, “Eliciting mixed emotions: a meta-analysis comparing models, types, and measures,” Frontiers in psychology , vol. 6, p. 428, 2015
2015
Cited alongside, same era.
H. Cao, R. Verma, and A. Nenkova, “Speaker-sensitive emotion recognition via ranking: Studies on acted and spontaneous speech,” Computer speech & language , vol. 29, no. 1, pp. 186–202, 2015
2015
Cited alongside, same era.
A. Rabiee, T.-H. Kim, and S.-Y. Lee, “Adjusting pleasure-arousal-dominance for continuous emotional text-to-speech synthesizer,” in INTERSPEECH 2019 . INTERSPEECH 2019, 2019
2019
Later among the works it cites.
V. Klimkov, S. Ronanki, J. Rohnke, and T. Drugman, “Fine-Grained Robust Prosody Transfer for Single-Speaker Neural Text-To-Speech,” in Proc. Interspeech 2019 , 2019, pp. 4440–4444
2019
Later among the works it cites.
Y. Lee and T. Kim, “Robust and fine-grained prosody control of end-to-end speech synthesis,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 5911–5915
2019
Later among the works it cites.
Y.-J. Zhang, S. Pan, L. He, and Z.-H. Ling, “Learning latent representations for style control and transfer in end-to-end speech synthesis,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6945–6949
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. N. Yannakakis and H. P. Martinez, “Grounding truth via ordinal annotation,” in 2015 international conference on affective computing and intelligent interaction (ACII) . IEEE, 2015, pp. 574–580
2015
Cited alongside, same era.
D. Bahdanau, K. H. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in 3rd International Conference on Learning Representations, ICLR 2015 , 2015
2015
Cited alongside, same era.
Z. Zhang, C. Wang, B. Xiao, W. Zhou, and S. Liu, “Robust relative attributes for human action recognition,” Pattern Analysis and Applications , vol. 18, no. 1, pp. 157–171, 2015
2015
Cited alongside, same era.
H. Muthusamy, K. Polat, and S. Yaacob, “Improved emotion recognition using gaussian mixture model and extreme learning machine in speech and glottal signals,” Mathematical Problems in Engineering , vol. 2015, 2015
2015
Cited alongside, same era.
G. Koch, R. Zemel, R. Salakhutdinov et al. , “Siamese neural networks for one-shot image recognition,” in ICML deep learning workshop , vol. 2. Lille, 2015, p. 0
2015
Cited alongside, same era.
J. Crumpton and C. L. Bethel, “A survey of using vocal prosody to convey emotion in robot speech,” International Journal of Social Robotics , vol. 8, no. 2, pp. 271–285, 2016
2016
Cited alongside, same era.
O. Watts, G. E. Henter, T. Merritt, Z. Wu, and S. King, “From hmms to dnns: where do the improvements come from?” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2016, pp. 5505–5509
2016
Cited alongside, same era.
H. Ming, D. Huang, L. Xie, J. Wu, M. Dong, and H. Li, “Deep bidirectional lstm modeling of timbre and prosody for emotional voice conversion,” Interspeech 2016 , pp. 2453–2457, 2016
2016
Cited alongside, same era.
T. Kenter, V. Wan, C.-A. Chan, R. Clark, and J. Vit, “Chive: Varying prosody in speech synthesis with a linguistically driven dynamic hierarchical conditional variational network,” in International Conference on Machine Learning . PMLR, 2019, pp. 3331–3340
2019
Later among the works it cites.
M. Zhang, X. Wang, F. Fang, H. Li, and J. Yamagishi, “Joint training framework for text-to-speech and voice conversion using multi-source tacotron and wavenet,” Proc. Interspeech 2019 , pp. 1298–1302, 2019
2019
Later among the works it cites.
J.-X. Zhang, Z.-H. Ling, and L.-R. Dai, “Non-parallel sequence-to-sequence voice conversion with disentangled linguistic and speaker representations,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 540–552, 2019
2019
Later among the works it cites.
A. Polyak and L. Wolf, “Attention-based wavenet autoencoder for universal voice conversion,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6800–6804
2019
Later among the works it cites.
H.-T. Luong and J. Yamagishi, “Bootstrapping non-parallel voice conversion from speaker-adaptive text-to-speech,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 200–207
2019
Later among the works it cites.
B. McFee and G. R. Lanckriet, “Metric learning to rank,” in ICML , 2010
2019
Later among the works it cites.
B. Sisman, J. Yamagishi, S. King, and H. Li, “An overview of voice conversion and its challenges: From statistical modeling to deep learning,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
D. M. Schuller and B. W. Schuller, “A review on five recent and near-future developments in computational processing of emotion in the human voice,” Emotion Review , p. 1754073919898526, 2020
2020
Later among the works it cites.
G. Zhang, Y. Qin, and T. Lee, “Learning syllable-level discrete prosodic representation for expressive speech generation.” in INTERSPEECH , 2020, pp. 3426–3430
2020
Later among the works it cites.
G. Sun, Y. Zhang, R. J. Weiss, Y. Cao, H. Zen, and Y. Wu, “Fully-hierarchical fine-grained prosody modeling for interpretable speech synthesis,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6264–6268
2020
Later among the works it cites.
G. Sun, Y. Zhang, R. J. Weiss, Y. Cao, H. Zen, A. Rosenberg, B. Ramabhadran, and Y. Wu, “Generating diverse and natural text-to-speech samples using a quantized fine-grained vae and autoregressive prosody prior,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6699–6703
2020
Later among the works it cites.
K. Matsumoto, S. Hara, and M. Abe, “Controlling the Strength of Emotions in Speech-Like Emotional Sound Generated by WaveNet,” in Proc. Interspeech 2020 , 2020, pp. 3421–3425
2020
Later among the works it cites.
S.-Y. Um, S. Oh, K. Byun, I. Jang, C. Ahn, and H.-G. Kang, “Emotional speech synthesis with rich and granularized control,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7254–7258
2020
Later among the works it cites.
W.-C. Huang, T. Hayashi, Y.-C. Wu, H. Kameoka, and T. Toda, “Voice transformer network: Sequence-to-sequence voice conversion using transformer with text-to-speech pretraining,” Proc. Interspeech 2020 , pp. 4676–4680, 2020
2020
Later among the works it cites.
H.-T. Luong and J. Yamagishi, “Nautilus: a versatile voice cloning system,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 2967–2981, 2020
2020
Later among the works it cites.
J.-X. Zhang, Z.-H. Ling, and L.-R. Dai, “Recognition-Synthesis Based Non-Parallel Voice Conversion with Adversarial Learning,” in Proc. Interspeech 2020 , 2020, pp. 771–775
2020
Later among the works it cites.
U. Tiwari, M. Soni, R. Chakraborty, A. Panda, and S. K. Kopparapu, “Multi-conditioning and data augmentation using generative noise model for speech emotion recognition in noisy conditions,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7194–7198
2020
Later among the works it cites.
K. Zhou, B. Sisman, M. Zhang, and H. Li, “Converting Anyone’s Emotion: Towards Speaker-Independent Emotional Voice Conversion,” in Proc. Interspeech 2020 , 2020, pp. 3416–3420
2020
Later among the works it cites.
K. Zhou, B. Sisman, and H. Li, “Transforming Spectrum and Prosody for Emotional Voice Conversion with Non-Parallel Training Data,” in Proc. Odyssey 2020 The Speaker and Language Recognition Workshop , 2020, pp. 230–237
2020
Later among the works it cites.
A. Rosenberg and J. Hirschberg, “Prosodic aspects of the attractive voice,” in Voice Attractiveness . Springer, 2021, pp. 17–40
2021
Later among the works it cites.
2021
Later among the works it cites.
K. Inoue, S. Hara, M. Abe, N. Hojo, and Y. Ijima, “Model architectures to extrapolate emotional expressions in dnn-based text-to-speech,” Speech Communication , vol. 126, pp. 35–43, 2021
2021
Later among the works it cites.
J. B. Harvill, S.-G. Leem, M. Abdelwahab, R. Lotfian, and C. Busso, “Quantifying emotional similarity in speech,” IEEE Transactions on Affective Computing , 2021
2021
Later among the works it cites.
Y. Lei, S. Yang, and L. Xie, “Fine-grained emotion strength transfer, control and prediction for emotional speech synthesis,” in 2021 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2021, pp. 423–430
2021
Later among the works it cites.
2021
Later among the works it cites.
X. Cai, D. Dai, Z. Wu, X. Li, J. Li, and H. Meng, “Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 5734–5738
2021
Later among the works it cites.
T. Li, S. Yang, L. Xue, and L. Xie, “Controllable emotion transfer for end-to-end speech synthesis,” in 2021 12th International Symposium on Chinese Spoken Language Processing (ISCSLP) . IEEE, 2021, pp. 1–5
2021
Later among the works it cites.
X. Li, C. Song, J. Li, Z. Wu, J. Jia, and H. Meng, “Towards Multi-Scale Style Control for Expressive Speech Synthesis,” in Proc. Interspeech 2021 , 2021, pp. 4673–4677
2021
Later among the works it cites.
K. Zhou, B. Sisman, and H. Li, “Limited Data Emotional Voice Conversion Leveraging Text-to-Speech: Two-Stage Sequence-to-Sequence Training,” in Proc. Interspeech 2021 , 2021, pp. 811–815
2021
Later among the works it cites.
H. Choi and M. Hahn, “Sequence-to-sequence emotional voice conversion with strength control,” IEEE Access , vol. 9, pp. 42 674–42 687, 2021
2021
Later among the works it cites.
D. Tan and T. Lee, “Fine-Grained Style Modeling, Transfer and Prediction in Text-to-Speech Synthesis via Phone-Level Content-Style Disentanglement,” in Proc. Interspeech 2021 , 2021, pp. 4683–4687
2021
Later among the works it cites.
M. Zhang, Y. Zhou, L. Zhao, and H. Li, “Transfer learning from speech synthesis to voice conversion with non-parallel training data,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 1290–1302, 2021
2021
Later among the works it cites.
K. Zhou, B. Sisman, R. Liu, and H. Li, “Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 920–924
2021
Later among the works it cites.
B. J. Abbaschian, D. Sierra-Sosa, and A. Elmaghraby, “Deep learning techniques for speech emotion recognition, from databases to models,” Sensors , vol. 21, no. 4, p. 1249, 2021
2021
Later among the works it cites.
——, “Vaw-gan for disentanglement and recomposition of emotional elements in speech,” in 2021 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2021, pp. 415–422
2021
Later among the works it cites.
K. Zhou, B. Sisman, R. Rana, B. W. Schuller, and H. Li, “Emotion intensity and its control for emotional voice conversion,” IEEE Transactions on Affective Computing , 2022
2022
Closest in time.
Y. Lei, S. Yang, X. Wang, and L. Xie, “Msemotts: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2022
2022
Closest in time.
T. Cornille, F. Wang, and J. Bekker, “Interactive multi-level prosody control for expressive speech synthesis,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 8312–8316
2022
Closest in time.
C.-B. Im, S.-H. Lee, S.-B. Kim, and S.-W. Lee, “Emoq-tts: Emotion intensity quantization for fine-grained controllable emotional text-to-speech,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 6317–6321
2022
Closest in time.
——, “Emotional voice conversion: Theory, databases and esd,” Speech Communication , vol. 137, pp. 1–18, 2022
2022
Closest in time.