Fetching the paper…
Reading the bibliography…
In this paper, we explore the use of pre-trained language models to learn sentiment information of written texts for speech sentiment analysis.
C. Cieri, D. Miller, and K. Walker, “The Fisher corpus: a resource for the next generations of speech-to-text.” in LREC , vol. 4, 2004, pp. 69–71
2004
Earlier work this paper cites.
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “IEMOCAP: Interactive emotional dyadic motion capture database,” Language resources and evaluation , vol. 42, no. 4, pp. 335–359, 2008
2008
Earlier work this paper cites.
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Ng, and C. Potts, “Recursive deep models for semantic compositionality over a sentiment treebank,” in EMNLP , 2013, pp. 1631–1642
2013
Earlier work this paper cites.
S. Mohammad, “A practical guide to sentiment annotation: Challenges and solutions,” in Proceedings of the 7th workshop on computational approaches to subjectivity, sentiment and social media analysis , 2016, pp. 174–179
2016
Earlier work this paper cites.
S. Mirsamadi, E. Barsoum, and C. Zhang, “Automatic speech emotion recognition using recurrent neural networks with local attention,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 2227–2231
2017
Earlier work this paper cites.
S. Kim, T. Hori, and S. Watanabe, “Joint CTC-attention based end-to-end speech recognition using multi-task learning,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 4835–4839
2017
Earlier work this paper cites.
P. Li, Y. Song, I. McLoughlin, W. Guo, and L. Dai, “An attention pooling based representation learning method for speech emotion recognition,” in Interspeech , 2018, pp. 3087–3091
2018
Earlier work this paper cites.
P. Tzirakis, J. Zhang, and B. W. Schuller, “End-to-end speech emotion recognition using deep neural networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 5089–5093
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. B. Zadeh, P. P. Liang, S. Poria, E. Cambria, and L.-P. Morency, “Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2018, pp. 2236–2246
2018
Cited alongside, same era.
J. Cho, R. Pappagari, P. Kulkarni, J. Villalba, Y. Carmiel, and N. Dehak, “Deep neural networks for emotion recognition combining audio and transcripts,” in Interspeech , 2018, pp. 247–251
2018
Cited alongside, same era.
Y.-P. Chen, R. Price, and S. Bangalore, “Spoken language understanding without speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 6189–6193
2018
Cited alongside, same era.
P. Haghani, A. Narayanan, M. Bacchiani, G. Chuang, N. Gaur, P. Moreno, R. Prabhavalkar, Z. Qu, and A. Waters, “From audio to semantics: Approaches to end-to-end spoken language understanding,” in IEEE Spoken Language Technology Workshop (SLT) , 2018, pp. 720–726
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
R. Li, Z. Wu, J. Jia, S. Zhao, and H. Meng, “Dilated residual network with multi-head self-attention for speech emotion recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 6675–6679
2019
Cited alongside, same era.
X. Wu, S. Liu, Y. Cao, X. Li, J. Yu, D. Dai, X. Ma, S. Hu, Z. Wu, X. Liu et al. , “Speech emotion recognition using capsule networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 6695–6699
2019
Cited alongside, same era.
Y. Xie, R. Liang, Z. Liang, C. Huang, C. Zou, and B. Schuller, “Speech emotion classification using attention-based LSTM,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 11, pp. 1675–1685, 2019
2019
Cited alongside, same era.
E. Kim and J. W. Shin, “DNN-based emotion recognition based on bottleneck acoustic features and lexical features,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 6720–6724
2019
Cited alongside, same era.
L. Lugosch, M. Ravanelli, P. Ignoto, V. S. Tomar, and Y. Bengio, “Speech model pre-training for end-to-end spoken language understanding,” in Interspeech , 2019, pp. 814–818
2019
Cited alongside, same era.
Z. Lu, L. Cao, Y. Zhang, C.-C. Chiu, and J. Fan, “Speech sentiment analysis via pre-trained features from end-to-end ASR models,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 7149–7153
2020
Later among the works it cites.
E. Chen, Z. Lu, H. Xu, L. Cao, Y. Zhang, and J. Fan, “A large scale speech sentiment corpus,” in Proceedings of The 12th Language Resources and Evaluation Conference , 2020, pp. 6549–6555
2020
Later among the works it cites.
S. Siriwardhana, A. Reis, R. Weerasekera, and S. Nanayakkara, “Jointly fine-tuning “BERT-like” self supervised models to improve multimodal speech emotion recognition,” in Interspeech , 2020, pp. 3755–3759
2020
Later among the works it cites.
2020
Later among the works it cites.