Fetching the paper…
Reading the bibliography…
Speech emotion recognition (SER) is the task of recognising human's emotional states from speech.
“Survey on speech emotion recognition: Features, classification schemes, and databases,”
Moataz El Ayadi, Mohamed S Kamel, and Fakhri Karray, · 2011
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Speech emotion recognition: Two decades in a nutshell, benchmarks, and ongoing trends,”
Björn W Schuller, · 2018
Earlier work this paper cites.
“Speech emotion recognition using deep learning techniques: A review,”
Ruhul Amin Khalil, Edward Jones, Mohammad Inayatullah Babar, Tariqullah Jan, Mohammad Haseeb Zafar, and Thamer Alhussain, · 2019
Earlier work this paper cites.
“wav2vec: Unsupervised pre-training for speech recognition,”
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli, · 2019
Earlier work this paper cites.
“Be your own teacher: Improve the performance of convolutional neural networks via self distillation,”
Linfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen, Chenglong Bao, and Kaisheng Ma, · 2019
Earlier work this paper cites.
“DEMoS: An Italian emotional speech corpus,”
Emilia Parada-Cabaleiro, Giovanni Costantini, Anton Batliner, Maximilian Schmitt, and Björn Schuller, · 2019
Earlier work this paper cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Earlier work this paper cites.
“vq-wav2vec: Self-supervised learning of discrete speech representations,”
Alexei Baevski, Steffen Schneider, and Michael Auli, · 2020
Earlier work this paper cites.
“Self-distillation amplifies regularization in hilbert space,”
Hossein Mobahi, Mehrdad Farajtabar, and Peter Bartlett, · 2020
Cited alongside, same era.
“Self-supervised label augmentation via input transformations,”
Hankook Lee, Sung Ju Hwang, and Jinwoo Shin, · 2020
Cited alongside, same era.
“Generating and protecting against adversarial attacks for deep speech-based emotion recognition models,”
Zhao Ren, Alice Baird, Jing Han, Zixing Zhang, and Björn Schuller, · 2020
Cited alongside, same era.
“Enhancing transferability of black-box adversarial attacks via lifelong learning for speech emotion recognition models,”
Zhao Ren, Jing Han, Nicholas Cummins, and Björn Schuller, · 2020
Cited alongside, same era.
“Self-distillation: Towards efficient and compact neural networks,”
Linfeng Zhang, Chenglong Bao, and Kaisheng Ma, · 2021
Cited alongside, same era.
“Arabic speech emotion recognition employing wav2vec2. 0 and hubert based on baved dataset,”
Omar Mohamed and Salah A Aly, · 2021
Later among the works it cites.
“Deep model compression and architecture optimization for embedded systems: A survey,”
Anthony Berthelier, Thierry Chateau, Stefan Duffner, Christophe Garcia, and Christophe Blanc, · 2021
Later among the works it cites.
“Towards model compression for deep learning based speech enhancement,”
Ke Tan and DeLiang Wang, · 2021
Later among the works it cites.
“Distilhubert: Speech representation learning by layer-wise distillation of hidden-unit bert,”
Heng-Jui Chang, Shu-wen Yang, and Hung-yi Lee, · 2022
Closest in time.
“Ssast: Self-supervised audio spectrogram transformer,”
Yuan Gong, Cheng-I Lai, Yu-An Chung, and James Glass, · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Prabhav Singh, Ridam Srivastava, KPS Rana, and Vineet Kumar, · 2021
Cited alongside, same era.
“A comprehensive review of speech emotion recognition systems,”
Taiba Majid Wani, Teddy Surya Gunawan, Syed Asif Ahmad Qadri, Mira Kartiwi, and Eliathamby Ambikairajah, · 2021
Cited alongside, same era.
“Emotion recognition from speech using wav2vec 2.0 embeddings,”
Leonardo Pepino, Pablo Riera, and Luciana Ferrer, · 2021
Cited alongside, same era.
“Exploring wav2vec 2.0 fine-tuning for improved speech emotion recognition,”
Li-Wei Chen and Alexander Rudnicky, · 2021
Cited alongside, same era.
Johannes Wagner, Andreas Triantafyllopoulos, Hagen Wierstorf, Maximilian Schmitt, Florian Eyben, and Björn W Schuller, · 2022
Closest in time.
“Revisiting self-distillation,”
Minh Pham, Minsu Cho, Ameya Joshi, and Chinmay Hegde, · 2022
Closest in time.
“Fine-tuning wav2vec2 for speaker recognition,”
Nik Vaessen and David A Van Leeuwen, · 2022
Closest in time.
“Real-time end-to-end speech emotion recognition with cross-domain adaptation,”
Konlakorn Wongpatikaseree, Sattaya Singkul, Narit Hnoohom, and Sumeth Yuenyong, · 2022
Closest in time.