Fetching the paper…
Reading the bibliography…
Speech emotion recognition (SER) has made significant strides with the advent of powerful self-supervised learning (SSL) models.
F. Burkhardt, A. Paeschke, M. Rolfes, W. F. Sendlmeier, B. Weiss et al. , “A database of german emotional speech.” in Interspeech , vol. 5, 2005, pp. 1517–1520
2005
Earlier work this paper cites.
O. Martin, I. Kotsia, B. Macq, and I. Pitas, “The enterface’05 audio-visual emotion database,” in 22nd international conference on data engineering workshops (ICDEW’06) . IEEE, 2006, pp. 8–8
2006
Earlier work this paper cites.
T. Wu, Y. Yang, Z. Wu, and D. Li, “Masc: A speech corpus in mandarin for emotion analysis and affective speaker recognition,” in 2006 IEEE Odyssey-the speaker and language recognition workshop . IEEE, 2006, pp. 1–5
2006
Earlier work this paper cites.
J. Tao, F. Liu, M. Zhang, and H. Jia, “Design of speech corpus for Mandarin text to speech,” in The Blizzard Challenge Workshop , 2008
2008
Earlier work this paper cites.
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “IEMOCAP: Interactive emotional dyadic motion capture database,” in Proc. LREC , 2008
2008
Earlier work this paper cites.
R. Altrov and H. Pajupuu, “Estonian emotional speech corpus: culture and age in selecting corpus testers,” in Human Language Technologies–The Baltic Perspective . IOS Press, 2010, pp. 25–32
2010
Earlier work this paper cites.
G. Costantini, I. Iaderola, A. Paoloni, and M. Todisco, “EMOVO corpus: an Italian emotional speech database,” in Proc. LREC , 2014
2014
Earlier work this paper cites.
S. Zhalehpour, O. Onder, Z. Akhtar, and C. E. Erdem, “Baum-1: A spontaneous audio-visual face database of affective and mental states,” IEEE Transactions on Affective Computing , vol. 8, no. 3, pp. 300–313, 2016
2016
Earlier work this paper cites.
S. Latif, A. Qayyum, M. Usman, and J. Qadir, “Cross lingual speech emotion recognition: Urdu vs. western languages,” in 2018 International conference on frontiers of information technology (FIT) . IEEE, 2018, pp. 88–93
2018
Earlier work this paper cites.
N. Vryzas, R. Kotsakis, A. Liatsou, C. A. Dimoulas, and G. Kalliris, “Speech emotion recognition for performance interaction,” Journal of the Audio Engineering Society , vol. 66, no. 6, pp. 457–467, 2018
2018
Earlier work this paper cites.
P. Gournay, O. Lahaie, and R. Lefebvre, “A Canadian French emotional speech dataset,” in Proc. ACM Multimedia , 2018
2018
Earlier work this paper cites.
S. Poria, D. Hazarika, N. Majumder, G. Naik, E. Cambria, and R. Mihalcea, “MELD: A multimodal multi-party dataset for emotion recognition in conversations,” in Proc. ACL , 2019
2019
Earlier work this paper cites.
O. Mohamad Nezami, P. Jamshid Lou, and M. Karami, “ShEMO: a large-scale validated database for Persian speech emotion detection,” in Proc. LREC , 2019
2019
Earlier work this paper cites.
R. Lotfian and C. Busso, “Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings,” IEEE Transactions on Affective Computing , vol. 10, pp. 471–483, 10 2019
2019
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in neural information processing systems , vol. 33, pp. 12 449–12 460, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
E. Parada-Cabaleiro, G. Costantini, A. Batliner, M. Schmitt, and B. W. Schuller, “Demos: An italian emotional speech corpus: Elicitation methods, machine learning, and perception,” Language Resources and Evaluation , vol. 54, no. 2, pp. 341–383, 2020
2020
Cited alongside, same era.
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al. , “Wavlm: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1505–1518, 2022
2022
Later among the works it cites.
N. Scheidwasser-Clow, M. Kegler, P. Beckmann, and M. Cernak, “Serab: A multi-lingual benchmark for speech emotion recognition,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 7697–7701
2022
Later among the works it cites.
2023
Later among the works it cites.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in International Conference on Machine Learning . PMLR, 2023, pp. 28 492–28 518
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Wang, Q. Wu, L. Song, Z. Yang, W. Wu, C. Qian, R. He, Y. Qiao, and C. C. Loy, “MEAD: A large-scale audio-visual dataset for emotional talking-face generation,” in Proc. ECCV , 2020
2020
Cited alongside, same era.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3451–3460, 2021
2021
Cited alongside, same era.
L. Pepino, P. Riera, and L. Ferrer, “Emotion recognition from speech using wav2vec 2.0 embeddings,” in Proc. Interspeech , 2021
2021
Cited alongside, same era.
M. M. Duville, L. M. Alonso-Valerdi, and D. I. Ibarra-Zarate, “The mexican emotional speech database (mesd): elaboration and assessment based on machine learning,” in 2021 43rd Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC) . IEEE, 2021, pp. 1644–1647
2021
Cited alongside, same era.
T. Müller and D. Kreutz, “Thorsten-voice dataset 2021.02,” Sep. 2021, Please use it to make the world a better place for whole humankind. [Online]. Available: https://doi.org/10.5281/zenodo.5525342
2021
Cited alongside, same era.
S. Sultana, M. S. Rahman, M. R. Selim, and M. Z. Iqbal, “SUST Bangla emotional speech corpus (SUBESCO): An audio-only emotional speech corpus for Bangla,” in Proc. PloS One , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Y.-A. Chung, Y. Zhang, W. Han, C.-C. Chiu, J. Qin, R. Pang, and Y. Wu, “W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,” in 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2021, pp. 244–250
2021
Cited alongside, same era.
2023
Later among the works it cites.
A. Baevski, A. Babu, W.-N. Hsu, and M. Auli, “Efficient self-supervised learning with contextualized target representations for vision, speech and language,” in International Conference on Machine Learning . PMLR, 2023, pp. 1416–1429
2023
Later among the works it cites.
Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
B. Elizalde, S. Deshmukh, M. Al Ismail, and H. Wang, “Clap learning audio concepts from natural language supervision,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
M. Agarla, S. Bianco, L. Celona, P. Napoletano, A. Petrovsky, F. Piccoli, R. Schettini, and I. Shanin, “Semi-supervised cross-lingual speech emotion recognition,” Expert Systems with Applications , vol. 237, p. 121368, 2024
2024
Closest in time.