Fetching the paper…
Reading the bibliography…
Human speech can be characterized by different components, including semantic content, speaker identity and prosodic information.
J.-A. Bachorowski, “Vocal expression and perception of emotion,”
1999
Earlier work this paper cites.
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “IEMOCAP: Interactive emotional dyadic motion capture database,”
2008
Earlier work this paper cites.
L. Van der Maaten and G. Hinton, “Visualizing data using t-sne.”
2008
Earlier work this paper cites.
2010
Earlier work this paper cites.
M. D. Pell, A. Jaywant, L. Monetta, and S. A. Kotz, “Emotional speech processing: Disentangling the effects of prosody and semantic cues,”
2011
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “LibriSpeech: an ASR corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
V. Peddinti, D. Povey, and S. Khudanpur, “A time delay neural network architecture for efficient modeling of long temporal contexts,” in
2015
Earlier work this paper cites.
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,”
2015
Earlier work this paper cites.
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer, “Scheduled sampling for sequence prediction with recurrent neural networks,” in
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in
2015
Earlier work this paper cites.
Z. Luo, T. Takiguchi, and Y. Ariki, “Emotional voice conversion using deep neural networks with MCC and F0 features,” in
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
C. Busso, S. Parthasarathy, A. Burmania, M. AbdelWahab, N. Sadoughi, and E. M. Provost, “Msp-improv: An acted corpus of dyadic interactions to study emotion perception,”
2016
Earlier work this paper cites.
Z. Luo, J. Chen, T. Takiguchi, and Y. Ariki, “Emotional voice conversion with adaptive scales f0 based on wavelet transform using limited amount of emotional data,” in
2017
Earlier work this paper cites.
A. Schirmer and R. Adolphs, “Emotion perception from face, voice, and touch: comparisons and convergence,”
2017
Earlier work this paper cites.
E. Lakomkin, C. Weber, S. Magg, and S. Wermter, “Reusing neural speech representations for auditory emotion recognition,” in
2017
Earlier work this paper cites.
D. Le, Z. Aldeneh, and E. M. Provost, “Discretized continuous speech emotion recognition with multi-task deep recurrent neural network,” in
2017
Earlier work this paper cites.
B. Şişman, H. Li, and K. C. Tan, “Transformation of prosody in voice conversion,” in
2017
Earlier work this paper cites.
P. Li, Y. Song, I. V. McLoughlin, W. Guo, and L.-R. Dai, “An attention pooling based representation learning method for speech emotion recognition,” in
2018
Earlier work this paper cites.
Y. Wang, D. Stanton, Y. Zhang, R.-S. Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, and R. A. Saurous, “Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,” in
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in
2018
Earlier work this paper cites.
R. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. Weiss, R. Clark, and R. A. Saurous, “Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,” in
2018
Earlier work this paper cites.
J. S. Chung, A. Nagrani, and A. Zisserman, “VoxCeleb2: Deep speaker recognition,” in
2018
Earlier work this paper cites.
J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in
2018
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
P. Barros, N. Churamani, E. Lakomkin, H. Siqueira, A. Sutherland, and S. Wermter, “The omg-emotion behavior dataset,” in
2018
Cited alongside, same era.
S. R. Livingstone and F. A. Russo, “The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,”
2018
Cited alongside, same era.
Z. Zhao, Z. Bao, Z. Zhang, N. Cummins, H. Wang, and B. Schüller, “Attention-enhanced connectionist temporal classification for discrete speech emotion recognition,” in
2019
Cited alongside, same era.
Y.-J. Zhang, S. Pan, L. He, and Z.-H. Ling, “Learning latent representations for style control and transfer in end-to-end speech synthesis,” in
2019
Cited alongside, same era.
K. Qian, Y. Zhang, S. Chang, X. Yang, and M. Hasegawa-Johnson, “AutoVC: Zero-shot voice style transfer with only autoencoder loss,” in
2021
Later among the works it cites.
J. Weston, R. Lenain, U. Meepegama, and E. Fristed, “Learning de-identified representations of prosody from raw audio,” in
2021
Later among the works it cites.
2021
Later among the works it cites.
K. Zhou, B. Sisman, and H. Li, “Limited data emotional voice conversion leveraging text-to-speech: Two-stage sequence-to-sequence training,” in
2021
Later among the works it cites.
S. Yang, P. Chi, Y. Chuang, C. J. Lai, K. Lakhotia, Y. Y. Lin, A. T. Liu, J. Shi, X. Chang, G. Lin, T. Huang, W. Tseng, K. Lee, D. Liu, Z. Huang, S. Dong, S. Li, S. Watanabe, A. Mohamed, and H. Lee, “SUPERB: speech processing universal performance benchmark,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
N. Tishby, F. C. Pereira, and W. Bialek, “The information bottleneck method,” in
2019
Cited alongside, same era.
S.-H. Gao, M.-M. Cheng, K. Zhao, X.-Y. Zhang, M.-H. Yang, and P. Torr, “Res2Net: A new multi-scale backbone architecture,”
2019
Cited alongside, same era.
R. Lotfian and C. Busso, “Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings,”
2019
Cited alongside, same era.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A simple data augmentation method for automatic speech recognition,” in
2019
Cited alongside, same era.
L. Tarantino, P. N. Garner, A. Lazaridis
2019
Cited alongside, same era.
S. Schneider, A. Baevski, R. Collobert, and M. Auli, “wav2vec: Unsupervised pre-training for speech recognition,” in
2019
Cited alongside, same era.
A. L. Ruba and S. D. Pollak, “The development of emotion reasoning in infancy and early childhood,”
2020
Cited alongside, same era.
2021
Later among the works it cites.
W. Wu, C. Zhang, and P. C. Woodland, “Emotion recognition by fusing time synchronous and time asynchronous representations,” in
2021
Later among the works it cites.
Z. Zhao, Z. Bao, Z. Zhang, N. Cummins, S. Sun, H. Wang, J. Tao, and B. Schüller, “Self-attention transfer networks for speech emotion recognition,”
2021
Later among the works it cites.
Q. Cao, M. Hou, B. Chen, Z. Zhang, and G. Lu, “Hierarchical network based on the fusion of static and dynamic features for speech emotion recognition,” in
2021
Later among the works it cites.
X. Wu, Y. Cao, H. Lu, S. Liu, D. Wang, Z. Wu, X. Liu, and H. Meng, “Speech emotion recognition using sequential capsule networks,”
2021
Later among the works it cites.
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao
2022
Closest in time.
2022
Closest in time.
Y. Wang, Y. Xie, K. Zhao, H. Wang, and Q. Zhang, “Unsupervised quantized prosody representation for controllable speech synthesis,” in
2022
Closest in time.
B. T. Atmaja, A. Sasou, and M. Akagi, “Speech emotion and naturalness recognitions with multitask and single-task learnings,”
2022
Closest in time.
H.-C. Chou, W.-C. Lin, C.-C. Lee, and C. Busso, “Exploiting annotators’ typed description of emotion perception to maximize utilization of ratings for speech emotion recognition,” in
2022
Closest in time.
2022
Closest in time.
Z. Luo, S. Lin, R. Liu, J. Baba, Y. Yoshikawa, and H. Ishiguro, “Decoupling speaker-independent emotions for voice conversion via source-filter networks,”
2022
Closest in time.
L. Qu, C. Weber, and S. Wermter, “Lipsound2: Self-supervised pre-training for lip-to-speech reconstruction and lip reading,”
2022
Closest in time.
K. Zhou, B. Sisman, R. Liu, and H. Li, “Emotional voice conversion: Theory, databases and ESD,”
2022
Closest in time.
H. Zou, Y. Si, C. Chen, D. Rajan, and E. S. Chng, “Speech emotion recognition with co-attention based multi-level acoustic information,” in
2022
Closest in time.
A. Baevski, W. Hsu, Q. Xu, A. Babu, J. Gu, and M. Auli, “data2vec: A general framework for self-supervised learning in speech, vision and language,” in
2022
Closest in time.
C. H. Chan, K. Qian, Y. Zhang, and M. Hasegawa-Johnson, “Speechsplit2. 0: Unsupervised speech disentanglement for voice conversion without tuning autoencoder bottlenecks,” in
2022
Closest in time.
P. Yue, L. Qu, S. Zheng, and T. Li, “Multi-task learning for speech emotion and emotion intensity recognition,” in
2022
Closest in time.
A. Aftab, A. Morsali, S. Ghaemmaghami, and B. Champagne, “LIGHT-SERNET: A lightweight fully convolutional neural network for speech emotion recognition,” in
2022
Closest in time.
2022
Closest in time.
W. Chen, X. Xing, X. Xu, J. Pang, and L. Du, “SpeechFormer++: A Hierarchical Efficient Framework for Paralinguistic Speech Processing,”
2023
Closest in time.