Fetching the paper…
Reading the bibliography…
In this paper, we explored how to boost speech emotion recognition (SER) with the state-of-the-art speech pre-trained model (PTM), data2vec, text generation technique, GPT-4, and speech synthesis technique, Azure TTS.
“SSML: A speech synthesis markup language,”
Paul Taylor and Amy Isard, · 1997
Earlier work this paper cites.
“IEMOCAP: Interactive emotional dyadic motion capture database,”
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan, · 2008
Earlier work this paper cites.
“Audio augmentation for speech recognition,”
Tom Ko, Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Domain-adversarial training of neural networks,”
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor Lempitsky, · 2016
Earlier work this paper cites.
“mixup: Beyond empirical risk minimization,”
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz, · 2018
Earlier work this paper cites.
“On the robustness of speech emotion recognition for human-robot interaction with deep neural networks,”
Egor Lakomkin, Mohammad Ali Zamani, Cornelius Weber, Sven Magg, and Stefan Wermter, · 2018
Earlier work this paper cites.
“wav2vec: Unsupervised pre-training for speech recognition,”
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli, · 2019
Earlier work this paper cites.
“vq-wav2vec: Self-supervised learning of discrete speech representations,”
Alexei Baevski, Steffen Schneider, and Michael Auli, · 2019
Earlier work this paper cites.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le, · 2019
Earlier work this paper cites.
“Direct modelling of speech emotion from raw speech,”
Siddique Latif, Rajib Rana, Sara Khalifa, Raja Jurdak, and Julien Epps, · 2019
Earlier work this paper cites.
“Cyclegan-based emotion style transfer as data augmentation for speech emotion recognition.,”
Fang Bao, Michael Neumann, and Ngoc Thang Vu, · 2019
Earlier work this paper cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Earlier work this paper cites.
“Deep architecture enhancing robustness to noise, adversarial attacks, and cross-corpus setting for speech emotion recognition,”
Siddique Latif, Rajib Rana, Sara Khalifa, Raja Jurdak, and Björn W Schuller, · 2020
Earlier work this paper cites.
“x-vectors meet emotions: A study on dependencies between emotion and speaker recognition,”
Raghavendra Pappagari, Tianzi Wang, Jesus Villalba, Nanxin Chen, and Najim Dehak, · 2020
Earlier work this paper cites.
“Stargan for emotional speech conversion: Validated by data augmentation of end-to-end emotion recognition,”
Georgios Rizos, Alice Baird, Max Elliott, and Björn Schuller, · 2020
Earlier work this paper cites.
“Survey of deep representation learning for speech emotion recognition,”
Siddique Latif, Rajib Rana, Sara Khalifa, Raja Jurdak, Junaid Qadir, and Bjoern W Schuller, · 2021
Cited alongside, same era.
“SUPPERB: Speech processing universal performance benchmark,”
Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y Lin, Andy T Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, et al., · 2021
Cited alongside, same era.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Cited alongside, same era.
“Emotion recognition from speech using wav2vec 2.0 embeddings,”
Leonardo Pepino, Pablo Riera, and Luciana Ferrer, · 2021
Cited alongside, same era.
Yingzhi Wang, Abdelmoumene Boumadane, and Abdelwahab Heba, · 2021
“Effects of data augmentations on speech emotion recognition,”
Bagus Tris Atmaja and Akira Sasou, · 2022
Later among the works it cites.
“MT4SSL: Boosting self-supervised speech representation learning by integrating multiple targets,”
Ziyang Ma, Zhisheng Zheng, Changli Tang, Yujin Wang, and Xie Chen, · 2023
Closest in time.
“MMspeech: Multi-modal multi-task encoder-decoder pre-training for speech recognition,”
Xiaohuan Zhou, Jiaming Wang, Zeyu Cui, Shiliang Zhang, Zhijie Yan, Jingren Zhou, and Chang Zhou, · 2023
Closest in time.
“Pushing the limits of unsupervised unit discovery for SSL speech representation,”
Ziyang Ma, Zhisheng Zheng, Guanrou Yang, Yu Wang, Chao Zhang, and Xie Chen, · 2023
Closest in time.
“Reducing barriers to self-supervised learning: Hubert pre-training with academic compute,”
William Chen, Xuankai Chang, Yifan Peng, Zhaoheng Ni, Soumi Maiti, and Shinji Watanabe, · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio,”
Guoguo Chen, Shuzhou Chai, Guanbo Wang, Jiayu Du, et al., · 2021
Cited alongside, same era.
“Wavlm: Large-scale self-supervised pre-training for full stack speech processing,”
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al., · 2022
Cited alongside, same era.
“Data2vec: A general framework for self-supervised learning in speech, vision and language,”
Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli, · 2022
Cited alongside, same era.
“Tessp: text-enhanced self-supervised speech pre-training,”
Zhuoyuan Yao, Shuo Ren, Sanyuan Chen, Ziyang Ma, Pengcheng Guo, and Lei Xie, · 2022
Cited alongside, same era.
“Large-scale self-supervised speech representation learning for automatic speaker verification,”
Zhengyang Chen, Sanyuan Chen, Yu Wu, Yao Qian, Chengyi Wang, Shujie Liu, Yanmin Qian, and Michael Zeng, · 2022
Cited alongside, same era.
“Exploration of a self-supervised speech model: A study on emotional corpora,”
Yuanchao Li, Yumnah Mohamied, Peter Bell, and Catherine Lai, · 2022
Cited alongside, same era.
“Speech emotion recognition using self-supervised features,”
Edmilson Morais, Ron Hoory, Weizhong Zhu, Itai Gat, Matheus Damasceno, and Hagai Aronowitz, · 2022
Cited alongside, same era.
“Integrating emotion recognition with speech recognition and speaker diarisation for conversations,”
Wen Wu, Chao Zhang, and Philip C. Woodland, · 2023
Closest in time.
“Exploring wav2vec 2.0 fine tuning for improved speech emotion recognition,”
Li-Wei Chen and Alexander Rudnicky, · 2023
Closest in time.
“Dawn of the Transformer era in speech emotion recognition: closing the valence gap,”
Johannes Wagner, Andreas Triantafyllopoulos, Hagen Wierstorf, Maximilian Schmitt, Felix Burkhardt, Florian Eyben, and Björn W Schuller, · 2023
Closest in time.
“Towards paralinguistic-only speech representations for end-to-end speech emotion recognition,”
George Ioannides, Michael Owen, Andrew Fletcher, Viktor Rozgic, and Chao Wang, · 2023
Closest in time.
“Vesper: A compact and effective pretrained model for speech emotion recognition,”
Weidong Chen, Xiaofen Xing, Peihao Chen, and Xiangmin Xu, · 2023
Closest in time.
“A preliminary study on augmenting speech emotion recognition using a diffusion model,”
Ibrahim Malik, Siddique Latif, Raja Jurdak, and Björn Schuller, · 2023
Closest in time.
“Refashioning emotion recognition modelling: The advent of generalised large models,”
Zixing Zhang, Liyizhe Peng, Tao Pang, Jing Han, Huan Zhao, and Bjorn W Schuller, · 2023
Closest in time.
“GPT-4 technical report,” 2023
OpenAI, · 2023
Closest in time.
“Emodiff: Intensity controllable emotional text-to-speech with soft-label guidance,”
Yiwei Guo, Chenpeng Du, Xie Chen, and Kai Yu, · 2023
Closest in time.