Fetching the paper…
Reading the bibliography…
Speech emotion recognition (SER) is an important part of human-computer interaction, receiving extensive attention from both industry and academia.
F. Burkhardt, A. Paeschke, M. Rolfes, W. F. Sendlmeier, B. Weiss et al. , “A database of German emotional speech,” in Proc. Interspeech , 2005
2005
Earlier work this paper cites.
O. Martin, I. Kotsia, B. Macq, and I. Pitas, “The eNTERFACE’05 audio-visual emotion database,” in Proc. ICDE Workshop , 2006
2006
Earlier work this paper cites.
J. Tao, F. Liu, M. Zhang, and H. Jia, “Design of speech corpus for Mandarin text to speech,” in The Blizzard Challenge Workshop , 2008
2008
Earlier work this paper cites.
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “IEMOCAP: Interactive emotional dyadic motion capture database,” in Proc. LREC , 2008
2008
Earlier work this paper cites.
K. Dupuis and M. K. Pichora-Fuller, “Toronto emotional speech set (TESS) - younger talker_happy.” University of Toronto, 2010
2010
Earlier work this paper cites.
I. Steiner, M. Schröder, and A. Klepp, “The PAVOQUE corpus as a resource for analysis and synthesis of expressive speech,” in Proc. Phonetik & Phonologie , 2013
2013
Earlier work this paper cites.
H. Cao, D. G. Cooper, M. K. Keutmann, R. C. Gur, A. Nenkova, and R. Verma, “CREMA-D: Crowd-sourced emotional multimodal actors dataset,” in Proc. TAC , 2014
2014
Earlier work this paper cites.
G. Costantini, I. Iaderola, A. Paoloni, and M. Todisco, “EMOVO corpus: an Italian emotional speech database,” in Proc. LREC , 2014
2014
Earlier work this paper cites.
P. Jackson and S. Haq, “Surrey audio-visual expressed emotion (SAVEE) database.” University of Surrey, 2014
2014
Earlier work this paper cites.
N. Vryzas, R. Kotsakis, A. Liatsou, C. A. Dimoulas, and G. Kalliris, “Speech emotion recognition for performance interaction,” in Proc. AES , 2018
2018
Earlier work this paper cites.
P. Gournay, O. Lahaie, and R. Lefebvre, “A Canadian French emotional speech dataset,” in Proc. ACM Multimedia , 2018
2018
Earlier work this paper cites.
A. Adigwe, N. Tits, K. E. Haddad, S. Ostadabbas, and T. Dutoit, “The emotional voices database: Towards controlling the emotion dimension in voice generation systems,” in arXiv preprint , 2018
2018
Earlier work this paper cites.
J. James, L. Tian, and C. Watson, “An open source emotional speech corpus for human robot interaction applications,” Proc. Interspeech , 2018
2018
Earlier work this paper cites.
S. R. Livingstone and F. A. Russo, “The Ryerson audio-visual database of emotional speech and song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in North American English,” in Proc. PloS One , 2018
2018
Earlier work this paper cites.
S. Latif, A. Qayyum, M. Usman, and J. Qadir, “Cross lingual speech emotion recognition: Urdu vs. western languages,” in Proc. FIT , 2018
2018
Cited alongside, same era.
S. Poria, D. Hazarika, N. Majumder, G. Naik, E. Cambria, and R. Mihalcea, “MELD: A multimodal multi-party dataset for emotion recognition in conversations,” in Proc. ACL , 2019
2019
Cited alongside, same era.
O. Mohamad Nezami, P. Jamshid Lou, and M. Karami, “ShEMO: a large-scale validated database for Persian speech emotion detection,” in Proc. LREC , 2019
2019
Cited alongside, same era.
T. Landry Dejoli, Q. He, H. Yan, and Y. Li, “ASVP-ESD: A dataset and its benchmark for emotion recognition using both speech and non-speech utterances,” in Proc. GSJ , 2020
2020
Cited alongside, same era.
K. Wang, Q. Wu, L. Song, Z. Yang, W. Wu, C. Qian, R. He, Y. Qiao, and C. C. Loy, “MEAD: A large-scale audio-visual dataset for emotional talking-face generation,” in Proc. ECCV , 2020
W. Wu, M. Wu, and K. Yu, “Climate and weather: Inspecting depression detection via emotion recognition,” in Proc. ICASSP , 2022
2022
Later among the works it cites.
N. Scheidwasser-Clow, M. Kegler, P. Beckmann, and M. Cernak, “Serab: A multi-lingual benchmark for speech emotion recognition,” in Proc. ICASSP , 2022
2022
Later among the works it cites.
J. Zhao, T. Zhang, J. Hu, Y. Liu, Q. Jin, X. Wang, and H. Li, “M3ED: Multi-modal multi-scene multi-label emotional dialogue database,” in Proc. ACL , 2022
2022
Later among the works it cites.
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al. , “WavLM: Large-scale self-supervised pre-training for full stack speech processing,” in Proc. JSTSP , 2022
2022
Later among the works it cites.
A. Baevski, W.-N. Hsu, Q. Xu, A. Babu, J. Gu, and M. Auli, “data2vec: A general framework for self-supervised learning in speech, vision and language,” in Proc. ICML , 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
L. Martinez-Lucas, M. Abdelwahab, and C. Busso, “The MSP-conversation corpus,” in Proc. Interspeech , 2020
2020
Cited alongside, same era.
L. KERKENI, C. CLEDER, S.-R. Youssef, and K. RAOOF, “French emotional speech database-oréau,” 2020
2020
Cited alongside, same era.
M. Miesikowska and D. Swisulski, “Emotions in Polish speech recordings,” https://doi.org/10.34808/h46c-hb44 , 2020
2020
Cited alongside, same era.
S. F. Canpolat, Z. Ormanoğlu, and D. Zeyrek, “Turkish Emotion Voice Database (TurEV-DB),” in Proc. SLTU Workshop , 2020
2020
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Proc. NeurIPS , 2020
2020
Cited alongside, same era.
K. Zhou, B. Sisman, R. Liu, and H. Li, “Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,” in Proc. ICASSP , 2021
2021
Cited alongside, same era.
M. M. Duville, L. M. Alonso-Valerdi, and D. I. Ibarra-Zarate, “The Mexican emotional speech database (MESD): elaboration and assessment based on machine learning,” in Proc. EMBC , 2021
2021
Cited alongside, same era.
2022
Later among the works it cites.
N. Antoniou, A. Katsamanis, T. Giannakopoulos, and S. Narayanan, “Designing and evaluating speech emotion recognition systems: A reality check case study with IEMOCAP,” in Proc. ICASSP , 2023
2023
Later among the works it cites.
E. A. Retta, E. Almekhlafi, R. Sutcliffe, M. Mhamed, H. Ali, and J. Feng, “A new Amharic speech emotion dataset and classification benchmark,” in Proc. TALLIP , 2023
2023
Later among the works it cites.
K. A. Noriy, X. Yang, and J. J. Zhang, “EMNS/Imz/Corpus: An emotive single-speaker dataset for narrative storytelling in games, television and graphic novels,” in arXiv preprint , 2023
2023
Later among the works it cites.
F. Catania, “Speech emotion recognition in Italian using Wav2Vec 2,” in Authorea Preprints , 2023
2023
Later among the works it cites.
Z. Lian, H. Sun, L. Sun, K. Chen, M. Xu, K. Wang, K. Xu, Y. He, Y. Li, J. Zhao et al. , “MER 2023: Multi-label learning, modality robustness, and semi-supervised learning,” in Proc. ACM Multimedia , 2023
2023
Later among the works it cites.
A. Baevski, A. Babu, W.-N. Hsu, and M. Auli, “Efficient self-supervised learning with contextualized target representations for vision, speech and language,” in Proc. IMCL , 2023
2023
Later among the works it cites.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in Proc. ICML , 2023
2023
Later among the works it cites.
G.-T. Lin, C.-H. Chiang, and H.-y. Lee, “Advancing large language models to capture varied speaking styles and respond properly in spoken conversations,” in Proc. ACL , 2024
2024
Closest in time.
Z. Ma, Z. Zheng, J. Ye, J. Li, Z. Gao, S. Zhang, and X. Chen, “emotion2vec: Self-supervised pre-training for speech emotion representation,” in Proc. ACL Findings , 2024
2024
Closest in time.