Fetching the paper…
Reading the bibliography…
We propose EmoDistill, a novel speech emotion recognition (SER) framework that leverages cross-modal knowledge distillation during training to learn strong linguistic and prosodic representations of emotion from speech.
“Iemocap: Interactive emotional dyadic motion capture database,”
C. Busso, M. Bulut, C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, · 2008
Earlier work this paper cites.
“Opensmile: the munich versatile and fast open-source audio feature extractor,”
F. Eyben, M. Wöllmer, and B. Schuller, · 2010
Earlier work this paper cites.
“Learning salient features for speech emotion recognition using convolutional neural networks,”
Q. Mao, M. Dong, Z. Huang, and Y. Zhan, · 2014
Earlier work this paper cites.
“Distilling the knowledge in a neural network,”
G. Hinton, O. Vinyals, and J. Dean, · 2015
Earlier work this paper cites.
“The geneva minimalistic acoustic parameter set (gemaps) for voice research and affective computing,”
F. Eyben, K. R. Scherer, B. W. Schuller, J. Sundberg, E. André, C. Busso, L. Y. Devillers, J. Epps, P. Laukka, S. S. Narayanan, et al., · 2015
Earlier work this paper cites.
“Deep residual learning for image recognition,”
K. He, X. Zhang, S. Ren, and J. Sun, · 2016
Earlier work this paper cites.
“Investigation on joint representation learning for robust feature extraction in speech emotion recognition.,”
D. Luo, Y. Zou, and D. Huang, · 2018
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
J. Devlin, M. Chang, K. Lee, and K. Toutanova, · 2018
Earlier work this paper cites.
“Bimodal speech emotion recognition using pre-trained language models,”
V. Heusser, N. Freymuth, S. Constantin, and A. Waibel, · 2019
Earlier work this paper cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, · 2020
Earlier work this paper cites.
“Multimodal approach of speech emotion recognition using multi-level multi-head fusion attention-based recurrent neural network,”
N. Ho, H. Yang, S. Kim, and G. Lee, · 2020
Cited alongside, same era.
“Speech emotion recognition with local-global aware deep representation learning,”
J. Liu, Z. Liu, L. Wang, L. Guo, and J. Dang, · 2020
Cited alongside, same era.
“Speech sentiment analysis via pre-trained features from end-to-end asr models,”
Z. Lu, L. Cao, Y. Zhang, C. Chiu, and J. Fan, · 2020
Cited alongside, same era.
“HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,”
W.-N. Hsu, B. Bolte, Y. H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, · 2021
Cited alongside, same era.
“Emotion recognition from speech using wav2vec 2.0 embeddings,”
L. Pepino, P. Riera, and L. Ferrer, · 2021
Cited alongside, same era.
“Fusing asr outputs in joint training for speech emotion recognition,”
Y. Li, P. Bell, and C. Lai, · 2022
Later among the works it cites.
“Light-sernet: A lightweight fully convolutional neural network for speech emotion recognition,”
A. Aftab, A. Morsali, S. Ghaemmaghami, and B. Champagne, · 2022
Later among the works it cites.
“Speech emotion recognition with co-attention based multi-level acoustic information,”
H. Zou, Y. Si, C. Chen, D. Rajan, and E. S. Chng, · 2022
Later among the works it cites.
“Dawn of the transformer era in speech emotion recognition: closing the valence gap,”
J. Wagner, A. Triantafyllopoulos, H. Wierstorf, M. Schmitt, F. Burkhardt, F. Eyben, and B. W. Schuller, · 2023
Closest in time.
“Multistage linguistic conditioning of convolutional layers for speech emotion recognition,”
A. Triantafyllopoulos, U. Reichel, S. Liu, S. Huber, F. Eyben, and B. W. Schuller, · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Wang, A. Boumadane, and A. Heba, · 2021
Cited alongside, same era.
“Multimodal cross-and self-attention network for speech emotion recognition,”
L. Sun, B. Liu, J. Tao, and Z. Lian, · 2021
Cited alongside, same era.
“Hierarchical network based on the fusion of static and dynamic features for speech emotion recognition,”
Q. Cao, M. Hou, B. Chen, Z. Zhang, and G. Lu, · 2021
Cited alongside, same era.
“Speech emotion recognition using sequential capsule networks,”
X. Wu, Y. Cao, H. Lu, S. Liu, D. Wang, Z. Wu, X. Liu, and H. Meng, · 2021
Cited alongside, same era.
“Exploring attention mechanisms for multimodal emotion recognition in an emergency call center corpus,”
T. Deschamps-Berger, L. Lamel, and L. Devillers, · 2023
Closest in time.
“Audio representation learning by distilling video as privileged information,”
A. Hajavi and A. Etemad, · 2023
Closest in time.
“Fast yet effective speech emotion recognition with self-distillation,”
Z. Ren, T. T. Nguyen, Y. Chang, and B. W. Schuller, · 2023
Closest in time.
“Temporal modeling matters: A novel temporal emotional modeling approach for speech emotion recognition,”
J. Ye, X. Wen, Y. Wei, Y. Xu, K. Liu, and H. Shan, · 2023
Closest in time.