Fetching the paper…
Reading the bibliography…
This paper presents a paradigm that adapts general large-scale pretrained models (PTMs) to speech emotion recognition task.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 1877–1901
1901
Earlier work this paper cites.
O. Chetia Phukan, A. Balaji Buduru, and R. Sharma, “Transforming the embeddings: A lightweight technique for speech emotion recognition tasks,” in INTERSPEECH , 2023, pp. 1903–1907
1907
Earlier work this paper cites.
H. Robbins and S. Monro, “A stochastic approximation method,” The annals of mathematical statistics , pp. 400–407, 1951
1951
Earlier work this paper cites.
O.-W. Kwon, K. Chan, J. Hao, and T.-W. Lee, “Emotion recognition by speech signals,” in Interspeech , 2003, pp. 125–128
2003
Earlier work this paper cites.
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “IEMOCAP: Interactive emotional dyadic motion capture database,” Language Resources and Evaluation , vol. 42, no. 4, pp. 335–359, 2008
2008
Earlier work this paper cites.
X. Glorot, A. Bordes, and Y. Bengio, “Deep sparse rectifier neural networks,” in Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics , 2011, pp. 315–323
2011
Earlier work this paper cites.
M. El Ayadi, M. S. Kamel, and F. Karray, “Survey on speech emotion recognition: Features, classification schemes, and databases,” Pattern recognition , vol. 44, no. 3, pp. 572–587, 2011
2011
Earlier work this paper cites.
H. Cao, D. G. Cooper, M. K. Keutmann, R. C. Gur, A. Nenkova, and R. Verma, “CREMA-D: Crowd-sourced emotional multimodal actors dataset,” IEEE Transactions on Affective Computing , vol. 5, no. 4, pp. 377–390, 2014
2014
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proceedings of the 31st International Conference on Neural Information Processing Systems , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
S. Schneider, A. Baevski, R. Collobert, and M. Auli, “wav2vec: Unsupervised pre-training for speech recognition,” in Interspeech , 2019, pp. 3465–3469
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 12 449–12 460
2020
Earlier work this paper cites.
Q. Cao, H. Trivedi, A. Balasubramanian, and N. Balasubramanian, “DeFormer: Decomposing pre-trained transformers for faster question answering,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 2020, pp. 4487–4497
2020
Earlier work this paper cites.
R. Weng, H. Yu, S. Huang, S. Cheng, and W. Luo, “Acquiring knowledge from pre-trained model to neural machine translation,” in Proceedings of the AAAI conference on artificial intelligence , vol. 34, no. 05, 2020, pp. 9266–9273
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of Machine Learning Research , vol. 21, no. 1, pp. 5485–5551, 2020
2020
Earlier work this paper cites.
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , June 2020
2020
Earlier work this paper cites.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in International Conference on Machine Learning , 2020, pp. 1597–1607
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
W. Liu, P. Zhou, Z. Wang, Z. Zhao, H. Deng, and Q. Ju, “FastBERT: a self-distilling BERT with adaptive inference time,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 2020, pp. 6035–6044
2020
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning , 2021, pp. 8748–8763
2021
Cited alongside, same era.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3451–3460, 2021
2021
Cited alongside, same era.
B. T. Atmaja, A. Sasou, and M. Akagi, “Survey on bimodal speech emotion recognition from acoustic and linguistic information fusion,” Speech Communication , vol. 140, pp. 11–28, 2022
2022
Later among the works it cites.
S. Minaee, Y. Boykov, F. Porikli, A. Plaza, N. Kehtarnavaz, and D. Terzopoulos, “Image segmentation using deep learning: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 7, pp. 3523–3542, 2022
2022
Later among the works it cites.
J. Chen, H. Guo, K. Yi, B. Li, and M. Elhoseiny, “VisualGPT: Data-efficient adaptation of pretrained language models for image captioning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , June 2022, pp. 18 030–18 040
2022
Later among the works it cites.
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , “Flamingo: a visual language model for few-shot learning,” Advances in Neural Information Processing Systems , vol. 35, pp. 23 716–23 736, 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
D. Chen, Y. Li, M. Qiu, Z. Wang, B. Li, B. Ding, H. Deng, J. Huang, W. Lin, and J. Zhou, “Adabert: Task-adaptive BERT compression with differentiable neural architecture search,” in Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence , 2021, pp. 2463–2469
2021
Cited alongside, same era.
2021
Cited alongside, same era.
P. Tzirakis, A. Nguyen, S. Zafeiriou, and B. W. Schuller, “Speech emotion recognition using semantic information,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2021, pp. 6279–6283
2021
Cited alongside, same era.
U. Naseem, M. Khushi, V. Reddy, S. Rajendran, I. Razzak, and J. Kim, “Bioalbert: A simple and effective pre-trained language model for biomedical named entity recognition,” in International Joint Conference on Neural Networks , 2021, pp. 1–7
2021
Cited alongside, same era.
G. Chen, S. Ma, Y. Chen, L. Dong, D. Zhang, J. Pan, W. Wang, and F. Wei, “Zero-shot cross-lingual transfer of neural machine translation with multilingual pretrained encoders,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , 2021, pp. 15–26
2021
Cited alongside, same era.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations , 2021, pp. 1–21
2021
Cited alongside, same era.
P. Chen, D. Huang, D. He, X. Long, R. Zeng, S. Wen, M. Tan, and C. Gan, “Rspnet: Relative speed perception for unsupervised video representation learning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 35, no. 2, 2021, pp. 1045–1053
2021
Cited alongside, same era.
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” in International Conference on Machine Learning , 2021, pp. 8821–8831
2021
Cited alongside, same era.
2022
Later among the works it cites.
X.-Y. Zhao, Q.-S. Zhu, and J. Zhang, “Speech enhancement using self-supervised pre-trained model and vector quantization,” in Asia-Pacific Signal and Information Processing Association Annual Summit and Conference , 2022, pp. 330–334
2022
Later among the works it cites.
J.-L. Li and C.-C. Lee, “An enroll-to-verify approach for cross-task unseen emotion class recognition,” IEEE Transactions on Affective Computing , to be published, doi: 10.1109/TAFFC.2022.3183166
2022
Later among the works it cites.
W. Chen, X. Xing, X. Xu, J. Yang, and J. Pang, “Key-sparse transformer for multimodal speech emotion recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2022, pp. 6897–6901
2022
Later among the works it cites.
H. Zou, Y. Si, C. Chen, D. Rajan, and E. S. Chng, “Speech emotion recognition with co-attention based multi-level acoustic information,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2022, pp. 7367–7371
2022
Later among the works it cites.
2022
Later among the works it cites.
W. Chen, X. Xing, X. Xu, J. Pang, and L. Du, “SpeechFormer: A hierarchical efficient framework incorporating the characteristics of speech,” in Interspeech , 2022, pp. 346–350
2022
Later among the works it cites.
W. Fan, X. Xu, B. Cai, and X. Xing, “ISNet: Individual standardization network for speech emotion recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 1803–1814, 2022
2022
Later among the works it cites.
2023
Closest in time.
L.-W. Chen and A. Rudnicky, “Exploring wav2vec 2.0 fine tuning for improved speech emotion recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2023, pp. 1–5
2023
Closest in time.
X. Sun, P. Chen, L. Chen, C. Li, T. H. Li, M. Tan, and C. Gan, “Masked motion encoding for self-supervised video representation learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 2235–2245
2023
Closest in time.
2023
Closest in time.
Q.-S. Zhu, J. Zhang, Z.-Q. Zhang, and L.-R. Dai, “A joint speech enhancement and self-supervised representation learning framework for noise-robust speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp. 1927–1939, 2023
2023
Closest in time.
R. E. Zezario, S.-W. Fu, F. Chen, C.-S. Fuh, H.-M. Wang, and Y. Tsao, “Deep learning-based non-intrusive multi-objective speech assessment model with cross-domain features,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp. 54–70, 2023
2023
Closest in time.
W. Chen, X. Xing, X. Xu, J. Pang, and L. Du, “DST: Deformable speech transformer for emotion recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2023, pp. 1–5
2023
Closest in time.
S. Shen, F. Liu, and A. Zhou, “Mingling or misalignment? Temporal shift for speech emotion recognition with pre-trained representations,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2023, pp. 1–5
2023
Closest in time.
T. Gong, J. Belanich, K. Somandepalli, A. Nagrani, B. Eoff, and B. Jou, “LanSER: Language-model supported speech emotion recognition,” in INTERSPEECH , 2023, pp. 2408–2412
2023
Closest in time.
W. Chen, X. Xing, X. Xu, J. Pang, and L. Du, “SpeechFormer++: A hierarchical efficient framework for paralinguistic speech processing,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp. 775–788, 2023
2023
Closest in time.
X. Zhang and Y. Li, “A dual attention-based modality-collaborative fusion network for emotion recognition,” in INTERSPEECH , 2023, pp. 1468–1472
2023
Closest in time.
H. Zhao, B. Li, and Z. Zhang, “Speaker-aware cross-modal fusion architecture for conversational emotion recognition,” in INTERSPEECH , 2023, pp. 2718–2722
2023
Closest in time.
E. Jing, Y. Liu, Y. Chai, J. Sun, S. Samtani, Y. Jiang, and Y. Qian, “A deep interpretable representation learning method for speech emotion recognition,” Information Processing & Management , vol. 60, no. 6, p. 103501, 2023
2023
Closest in time.