Fetching the paper…
Reading the bibliography…
Multimodal emotion recognition plays a crucial role in enhancing user experience in human-computer interaction.
W. M. Wundt, An introduction to psychology . G. Allen, Limited, 1912
1912
Earlier work this paper cites.
M. M. Bradley and P. J. Lang, “Measuring emotion: the self-assessment manikin and the semantic differential,” Journal of Behavior Therapy and Experimental Psychiatry , vol. 25, no. 1, pp. 49–59, 1994
1994
Earlier work this paper cites.
L. F. Barrett, “Discrete emotions or dimensions? the role of valence focus and arousal focus,” Cognition & Emotion , vol. 12, no. 4, pp. 579–599, 1998
1998
Earlier work this paper cites.
R. W. Picard, Affective computing . MIT press, 2000
2000
Earlier work this paper cites.
Y.-I. Tian, T. Kanade, and J. F. Cohn, “Recognizing action units for facial expression analysis,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 23, no. 2, pp. 97–115, 2001
2001
Earlier work this paper cites.
Z. Zeng, M. Pantic, G. I. Roisman, and T. S. Huang, “A survey of affect recognition methods: audio, visual and spontaneous expressions,” in Proceedings of the 9th International Conference on Multimodal Interfaces , 2007, pp. 126–133
2007
Earlier work this paper cites.
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “Iemocap: Interactive emotional dyadic motion capture database,” Language Resources and Evaluation , vol. 42, pp. 335–359, 2008
2008
Earlier work this paper cites.
B. Schuller, S. Steidl, and A. Batliner, “The interspeech 2009 emotion challenge,” in Proceedings of the Interspeech , 2009, pp. 312–315
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2009, pp. 248–255
2009
Earlier work this paper cites.
B. Schuller, S. Steidl, A. Batliner, F. Burkhardt, L. Devillers, C. Müller, and S. Narayanan, “The interspeech 2010 paralinguistic challenge,” in Proceedings of the Interspeech , 2010, pp. 2794–2797
2010
Earlier work this paper cites.
P. K. Atrey, M. A. Hossain, A. El Saddik, and M. S. Kankanhalli, “Multimodal fusion for multimedia analysis: a survey,” Multimedia Systems , vol. 16, pp. 345–379, 2010
2010
Earlier work this paper cites.
S. Hamann, “Mapping discrete and dimensional emotions onto the brain: controversies and consensus,” Trends in Cognitive Sciences , vol. 16, no. 9, pp. 458–466, 2012
2012
Earlier work this paper cites.
I. J. Goodfellow, D. Erhan, P. L. Carrier, A. Courville, M. Mirza, B. Hamner, W. Cukierski, Y. Tang, D. Thaler, D.-H. Lee et al. , “Challenges in representation learning: A report on three machine learning contests,” in Proceedings of the 20th International Conference on Neural Information Processing (ICONIP) , 2013, pp. 117–124
2013
Earlier work this paper cites.
P. Chandrasekar, S. Chapaneri, and D. Jayaswal, “Automatic speech emotion recognition: A survey,” in 2014 International Conference on Circuits, Systems, Communication and Information Technology Applications (CSCITA) . IEEE, 2014, pp. 341–346
2014
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” Journal of Machine Learning Research , vol. 15, no. 1, pp. 1929–1958, 2014
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proceedings of the International Conference on Learning Representations , 2015, pp. 1–15
2015
Earlier work this paper cites.
S. E. Kahou, X. Bouthillier, P. Lamblin, C. Gulcehre, V. Michalski, K. Konda, S. Jean, P. Froumenty, Y. Dauphin, N. Boulanger-Lewandowski et al. , “Emonets: Multimodal deep learning approaches for emotion recognition in video,” Journal on Multimodal User Interfaces , vol. 10, pp. 99–111, 2016
2016
Earlier work this paper cites.
K. Weiss, T. M. Khoshgoftaar, and D. Wang, “A survey of transfer learning,” Journal of Big Data , vol. 3, no. 1, pp. 1–40, 2016
2016
Earlier work this paper cites.
G. Trigeorgis, F. Ringeval, R. Brueckner, E. Marchi, M. A. Nicolaou, B. Schuller, and S. Zafeiriou, “Adieu features? end-to-end speech emotion recognition using a deep convolutional recurrent network,” in Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2016, pp. 5200–5204
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
Y. Guo, L. Zhang, Y. Hu, X. He, and J. Gao, “Ms-celeb-1m: A dataset and benchmark for large-scale face recognition,” in Proceedings of the European Conference on Computer Vision , 2016, pp. 87–102
2016
Earlier work this paper cites.
T. Baltrušaitis, P. Robinson, and L.-P. Morency, “Openface: an open source facial behavior analysis toolkit,” in Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV) . IEEE, 2016, pp. 1–10
2016
Earlier work this paper cites.
S. Poria, E. Cambria, R. Bajpai, and A. Hussain, “A review of affective computing: From unimodal analysis to multimodal fusion,” Information Fusion , vol. 37, pp. 98–125, 2017
2017
Earlier work this paper cites.
S. Poria, E. Cambria, D. Hazarika, N. Majumder, A. Zadeh, and L.-P. Morency, “Context-dependent sentiment analysis in user-generated videos,” in Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics , vol. 1, 2017, pp. 873–883
2017
Earlier work this paper cites.
A. Zadeh, M. Chen, S. Poria, E. Cambria, and L.-P. Morency, “Tensor fusion network for multimodal sentiment analysis,” in Proceedings of the Conference on Empirical Methods in Natural Language Processing , 2017, pp. 1103–1114
2017
Earlier work this paper cites.
S. Li, W. Deng, and J. Du, “Reliable crowdsourcing and deep locality-preserving learning for expression recognition in the wild,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2017, pp. 2852–2861
2017
Earlier work this paper cites.
D. Hendrycks and K. Gimpel, “A baseline for detecting misclassified and out-of-distribution examples in neural networks,” in Proceedings of the International Conference on Learning Representations , 2017, pp. 1–12
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
T. Young, D. Hazarika, S. Poria, and E. Cambria, “Recent trends in deep learning based natural language processing,” IEEE Computational Intelligence Magazine , vol. 13, no. 3, pp. 55–75, 2018
2018
Earlier work this paper cites.
A. Zadeh, P. P. Liang, S. Poria, P. Vij, E. Cambria, and L.-P. Morency, “Multi-attention recurrent network for human communication comprehension,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2018, pp. 5642–5649
2018
Earlier work this paper cites.
A. B. Zadeh, P. P. Liang, S. Poria, E. Cambria, and L.-P. Morency, “Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2018, pp. 2236–2246
2018
Earlier work this paper cites.
D. Hazarika, S. Poria, R. Mihalcea, E. Cambria, and R. Zimmermann, “Icon: interactive conversational memory network for multimodal emotion detection,” in Proceedings of the Conference on Empirical Methods in Natural Language Processing , 2018, pp. 2594–2604
2018
Earlier work this paper cites.
J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2018, pp. 7132–7141
2018
Earlier work this paper cites.
Z. Liu, Y. Shen, V. B. Lakshminarasimhan, P. P. Liang, A. Zadeh, and L.-P. Morency, “Efficient low-rank multimodal fusion with modality-specific factors,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics , 2018, pp. 2247–2256
2018
Earlier work this paper cites.
A. Zadeh, P. P. Liang, N. Mazumder, S. Poria, E. Cambria, and L.-P. Morency, “Memory fusion network for multi-view sequential learning,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2018, pp. 5634–5641
2018
Earlier work this paper cites.
D. Hazarika, S. Poria, A. Zadeh, E. Cambria, L.-P. Morency, and R. Zimmermann, “Conversational memory network for emotion recognition in dyadic dialogue videos,” in Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2018, pp. 2122–2132
2018
Cited alongside, same era.
C.-C. Hsu, S.-Y. Chen, C.-C. Kuo, T.-H. K. Huang, and L.-W. Ku, “Emotionlines: An emotion corpus of multi-party conversations,” in Proceedings of the Eleventh International Conference on Language Resources and Evaluation , 2018, pp. 1597–1601
2018
Cited alongside, same era.
J. S. Chung, A. Nagrani, and A. Zisserman, “Voxceleb2: Deep speaker recognition,” in Proceedings of the Interspeech , 2018, pp. 1086–1090
2018
Cited alongside, same era.
N. Majumder, S. Poria, D. Hazarika, R. Mihalcea, A. Gelbukh, and E. Cambria, “Dialoguernn: An attentive rnn for emotion detection in conversations,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2019, pp. 6818–6825
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3451–3460, 2021
2021
Later among the works it cites.
Z. Zhao, Q. Liu, and S. Wang, “Learning deep global multi-scale and local attention features for facial expression recognition in the wild,” IEEE Transactions on Image Processing , vol. 30, pp. 6544–6556, 2021
2021
Later among the works it cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in Proceedings of the International Conference on Machine Learning . PMLR, 2021, pp. 8748–8763
2021
Later among the works it cites.
W. Han, H. Chen, and S. Poria, “Improving multimodal fusion with hierarchical mutual information maximization for multimodal sentiment analysis,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , 2021, pp. 9180–9192
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
S. Poria, D. Hazarika, N. Majumder, G. Naik, E. Cambria, and R. Mihalcea, “Meld: A multimodal multi-party dataset for emotion recognition in conversations,” in Proceedings of the 57th Conference of the Association for Computational Linguistics , 2019, pp. 527–536
2019
Cited alongside, same era.
S. Schneider, A. Baevski, R. Collobert, and M. Auli, “wav2vec: Unsupervised pre-training for speech recognition,” in Proceedings of the Interspeech , 2019, pp. 3465–3469
2019
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2019, pp. 4171–4186
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V. Le, “Xlnet: Generalized autoregressive pretraining for language understanding,” Proceedings of the Advances in Neural Information Processing Systems , pp. 5754–5764, 2019
2019
Cited alongside, same era.
H. Pham, P. P. Liang, T. Manzini, L.-P. Morency, and B. Póczos, “Found in translation: Learning robust joint representations by cyclic translations between modalities,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2019, pp. 6892–6899
2019
Cited alongside, same era.
Y.-H. H. Tsai, P. P. Liang, A. Zadeh, L.-P. Morency, and R. Salakhutdinov, “Learning factorized multimodal representations,” in Proceedings of the 7th International Conference on Learning Representations , 2019, pp. 1–20
2019
Cited alongside, same era.
Y.-H. H. Tsai, S. Bai, P. P. Liang, J. Z. Kolter, L.-P. Morency, and R. Salakhutdinov, “Multimodal transformer for unaligned multimodal language sequences,” in Proceedings of the 57th Conference of the Association for Computational Linguistics , 2019, pp. 6558–6569
2019
Cited alongside, same era.
2021
Later among the works it cites.
Z. Yao, D. Wu, X. Wang, B. Zhang, F. Yu, C. Yang, Z. Peng, X. Chen, L. Xie, and X. Lei, “Wenet: Production oriented streaming and non-streaming end-to-end speech recognition toolkit,” in Proceedings of the Interspeech , 2021, pp. 4054–4058
2021
Later among the works it cites.
2021
Later among the works it cites.
Z. Lian, B. Liu, and J. Tao, “Smin: Semi-supervised multi-modal interaction network for conversational emotion recognition,” IEEE Transactions on Affective Computing , 2022
2022
Later among the works it cites.
J. Zhao, T. Zhang, J. Hu, Y. Liu, Q. Jin, X. Wang, and H. Li, “M3ed: Multi-modal multi-scene multi-label emotional dialogue database,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics , 2022, pp. 5699–5710
2022
Later among the works it cites.
Y. Liu, Z. Yuan, H. Mao, Z. Liang, W. Yang, Y. Qiu, T. Cheng, X. Li, H. Xu, and K. Gao, “Make acoustic and visual cues matter: Ch-sims v2. 0 dataset and av-mixup consistent module,” in Proceedings of the 2022 International Conference on Multimodal Interaction , 2022, pp. 247–258
2022
Later among the works it cites.
Z. Tong, Y. Song, J. Wang, and L. Wang, “Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training,” in Proceedings of the Advances in Neural Information Processing Systems , 2022, pp. 10 078–10 093
2022
Later among the works it cites.
C. Zhang, Y. Cui, Z. Han, J. T. Zhou, H. Fu, and Q. Hu, “Deep partial multi-view learning,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 05, pp. 2402–2415, 2022
2022
Later among the works it cites.
A. Mohamed, H. Y. Lee, L. Borgholt, J. D. Havtorn, J. Edin, C. Igel, K. Kirchhoff, S. W. Li, K. Livescu, L. Maaløe et al. , “Self-supervised speech representation learning: A review,” IEEE Journal on Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1179–1210, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
N. Karim, M. N. Rizve, N. Rahnavard, A. Mian, and M. Shah, “Unicon: Combating label noise through uniform selection and contrastive learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 9676–9686
2022
Later among the works it cites.
A. Baevski, W.-N. Hsu, Q. Xu, A. Babu, J. Gu, and M. Auli, “Data2vec: A general framework for self-supervised learning in speech, vision and language,” in Proceedings of the International Conference on Machine Learning . PMLR, 2022, pp. 1298–1312
2022
Later among the works it cites.
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al. , “Wavlm: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1505–1518, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Z. Du, Y. Qian, X. Liu, M. Ding, J. Qiu, Z. Yang, and J. Tang, “Glm: General language model pretraining with autoregressive blank infilling,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2022, pp. 320–335
2022
Later among the works it cites.
S. Amiriparian, L. Christ, A. König, A. Cowen, E.-M. Meßner, E. Cambria, and B. W. Schuller, “Muse 2023 challenge: Multimodal prediction of mimicked emotions, cross-cultural humour, and personalised recognition of affects,” in Proceedings of the 31st ACM International Conference on Multimedia , 2023, pp. 9723–9725
2023
Later among the works it cites.
Z. Lian, L. Chen, L. Sun, B. Liu, and J. Tao, “Gcnet: Graph completion network for incomplete multimodal learning in conversation,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 07, pp. 8419–8432, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
M. C. Schiappa, Y. S. Rawat, and M. Shah, “Self-supervised learning for videos: A survey,” ACM Computing Surveys , vol. 55, no. 13s, pp. 1–37, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
A. Gandhi, K. Adhvaryu, S. Poria, E. Cambria, and A. Hussain, “Multimodal sentiment analysis: A systematic review of history, datasets, multimodal fusion methods, applications, challenges and future directions,” Information Fusion , vol. 91, pp. 424–444, 2023
2023
Later among the works it cites.
L. Sun, Z. Lian, B. Liu, and J. Tao, “Mae-dfer: Efficient masked autoencoder for self-supervised dynamic facial expression recognition,” in Proceedings of the 31st ACM International Conference on Multimedia , 2023, pp. 6110–6121
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in Proceedings of the International Conference on Machine Learning . PMLR, 2023, pp. 28 492–28 518
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing, “Vicuna: An open-source chatbot impressing gpt-4 with 90% chatgpt quality,” 2023. [Online]. Available: https://vicuna.lmsys.org
2023
Later among the works it cites.
J. Tow, “Stablelm alpha v2 models,” 2023. [Online]. Available: https://huggingface.co/stabilityai/stablelm-base-alpha-7b-v2
2023
Later among the works it cites.