Fetching the paper…
Reading the bibliography…
In this paper, we propose MMER, a novel Multimodal Multi-task learning approach for Speech Emotion Recognition.
W. Yang, S. Fukayama, P. Heracleous, and J. Ogata, “Exploiting Fine-tuning of Self-supervised Learning Models for Improving Bi-modal Sentiment Analysis and Emotion Recognition,” in Proc. Interspeech 2022 , 2022, pp. 1998–2002
2002
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proceedings of the 23rd international conference on Machine learning , 2006, pp. 369–376
2006
Earlier work this paper cites.
B. et al., “Iemocap: Interactive emotional dyadic motion capture database,” LREC 2008 , pp. 335–359
2008
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in IEEE ICASSP 2015 , pp. 5206–5210
2015
Earlier work this paper cites.
S. Poria, E. Cambria, D. Hazarika, N. Majumder, A. Zadeh, and L.-P. Morency, “Context-dependent sentiment analysis in user-generated videos,” in ACL 2017 , 2017, pp. 873–883
2017
Earlier work this paper cites.
V. et al., “Attention is all you need,” NeurIPS 2017 , vol. 30
2017
Earlier work this paper cites.
M. Sarma, P. Ghahremani, D. Povey, N. K. Goel, K. K. Sarma, and N. Dehak, “Emotion identification from raw speech signals using dnns.” in Interspeech 2018 , pp. 3097–3101
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
P. et al., “Multimodal sentiment analysis: Addressing key issues and setting up the baselines,” IEEE Intelligent Systems , vol. 33, no. 6, pp. 17–25, 2018
2018
Earlier work this paper cites.
W. Y. Choi, K. Y. Song, and C. W. Lee, “Convolutional attention networks for multimodal emotion recognition from speech and text data,” in 2018 Challenge-HML , 2018, pp. 28–34
2018
Earlier work this paper cites.
A. W. Yu, D. Dohan, M.-T. Luong, R. Zhao, K. Chen, M. Norouzi, and Q. V. Le, “Qanet: Combining local convolution with global self-attention for reading comprehension,” in ICLR 2018
2018
Earlier work this paper cites.
W. et al., “Speech emotion recognition using capsule networks,” in IEEE ICASSP 2019 , pp. 6695–6699
2019
Cited alongside, same era.
Y. Wang, Y. Shen, Z. Liu, P. P. Liang, A. Zadeh, and L.-P. Morency, “Words can shift: Dynamically adjusting word representations using nonverbal behaviors,” in AAAI 2019 , pp. 7216–7223
2019
Cited alongside, same era.
J. Sebastian, P. Pierucci et al. , “Fusion techniques for utterance-level emotion recognition combining speech and transcripts.” in Interspeech 2019 , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Y.-H. H. Tsai, S. Bai, P. P. Liang, J. Z. Kolter, L.-P. Morency, and R. Salakhutdinov, “Multimodal transformer for unaligned multimodal language sequences,” in ACL 2019 , p. 6558
J. Wang, M. Xue, R. Culhane, E. Diao, J. Ding, and V. Tarokh, “Speech emotion recognition with dual-sequence lstm architecture,” in IEEE ICASSP 2020 , 2020, pp. 6474–6478
2020
Later among the works it cites.
X. Cai, J. Yuan, R. Zheng, L. Huang, and K. Church, “Speech emotion recognition with multi-task learning,” in Interspeech 2021 , 2021
2021
Later among the works it cites.
M. R. Makiuchi, K. Uto, and K. Shinoda, “Multimodal emotion recognition with high-level speech and text features,” in IEEE ASRU 2021 , 2021, pp. 350–357
2021
Later among the works it cites.
A. Keesing, Y. S. Koh, and M. Witbrock, “Acoustic features and neural representations for categorical emotion recognition from speech,” in Interspeech 2021 , pp. 3415–3419
2021
Later among the works it cites.
Y. et al., “SUPERB: Speech Processing Universal PERformance Benchmark,” in Interspeech 2021 , 2021, pp. 1194–1198
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” NeurIPS 2020 , pp. 12 449–12 460
2020
Cited alongside, same era.
M. Sajjad, S. Kwon et al. , “Clustering-based speech emotion recognition by incorporating learned features and deep bilstm,” IEEE Access 2020 , vol. 8, pp. 79 861–79 875
2020
Cited alongside, same era.
Z. Lu, L. Cao, Y. Zhang, C.-C. Chiu, and J. Fan, “Speech sentiment analysis via pre-trained features from end-to-end asr models,” in IEEE ICASSP 2020 , pp. 7149–7153
2020
Cited alongside, same era.
P. Khosla, P. Teterwak, C. Wang, A. Sarna, Y. Tian, P. Isola, A. Maschinot, C. Liu, and D. Krishnan, “Supervised contrastive learning,” Advances in neural information processing systems , vol. 33, pp. 18 661–18 673, 2020
2020
Cited alongside, same era.
D. Krishna and A. Patil, “Multimodal emotion recognition using cross-modal attention and 1d convolutional neural networks.” in Interspeech 2020 , 2020, pp. 4243–4247
2020
Cited alongside, same era.
2020
Cited alongside, same era.
H. Li, M. Tu, J. Huang, S. Narayanan, and P. Georgiou, “Speaker-invariant affective representation learning via adversarial training,” in IEEE ICASSP 2020 , pp. 7144–7148
2020
Cited alongside, same era.
2021
Later among the works it cites.
X. Cai, J. Yuan, R. Zheng, L. Huang, and K. Church, “Speech Emotion Recognition with Multi-Task Learning,” in Interspeech 2021 , 2021, pp. 4508–4512
2021
Later among the works it cites.
W. Wu, C. Zhang, and P. C. Woodland, “Emotion recognition by fusing time synchronous and time asynchronous representations,” in IEEE ICASSP 2021 , 2021, pp. 6269–6273
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2022
Closest in time.
2022
Closest in time.
E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. Gölge, and M. A. Ponti, “Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,” in ICML 2022 , pp. 2709–2720
2022
Closest in time.