Fetching the paper…
Reading the bibliography…
Multimodal sentiment analysis aims to effectively integrate information from various sources to infer sentiment, where in many cases there are no annotations for unimodal labels.
D. Olson, “From utterance to text: The bias of language in speech and writing,” Harvard Educational Review , vol. 47, no. 3, pp. 257–281, 1977
1977
Earlier work this paper cites.
2008
Earlier work this paper cites.
G. Degottex, J. Kane, T. Drugman, T. Raitio, and S. Scherer, “Covarep: A collaborative voice analysis repository for speech technologies,” in ICASSP , 2014, pp. 960–964
2014
Earlier work this paper cites.
S. Sukhbaatar, J. Bruna, M. Paluri, L. Bourdev, and R. Fergus, “Training convolutional networks with noisy labels,” in 3rd International Conference on Learning Representations, ICLR 2015 , 2015
2015
Earlier work this paper cites.
B. McFee, C. Raffel, D. Liang, D. P. Ellis, M. McVicar, E. Battenberg, and O. Nieto, “librosa: Audio and music signal analysis in python,” in Proceedings of the 14th python in science conference , vol. 8, 2015, pp. 18–25
2015
Earlier work this paper cites.
A. Zadeh, R. Zellers, E. Pincus, and L. P. Morency, “Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages,” IEEE Intelligent Systems , vol. 31, no. 6, pp. 82–88, 11 2016
2016
Earlier work this paper cites.
K. Zhang, Z. Zhang, Z. Li, and Y. Qiao, “Joint face detection and alignment using multitask cascaded convolutional networks,” IEEE signal processing letters , vol. 23, no. 10, pp. 1499–1503, 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
S. Poria, E. Cambria, R. Bajpai, and A. Hussain, “A review of affective computing: From unimodal analysis to multimodal fusion,” Information Fusion , vol. 37, pp. 98–125, 2017
2017
Earlier work this paper cites.
C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,” ICML , pp. 1126–1135, 2017
2017
Earlier work this paper cites.
A. Zadeh, M. Chen, S. Poria, E. Cambria, and L. P. Morency, “Tensor fusion network for multimodal sentiment analysis,” in EMNLP , 2017, pp. 1114–1125
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
A. Zadeh, P. P. Liang, N. Mazumder, S. Poria, E. Cambria, and L. P. Morency, “Memory fusion network for multi-view sequential learning,” in AAAI , 2018, pp. 5634–5641
2018
Earlier work this paper cites.
A. Zadeh, P. P. Liang, S. Poria, P. Vij, E. Cambria, and L. P. Morency, “Multi-attention recurrent network for human communication comprehension,” in AAAI , 2018, pp. 5642–5649
2018
Earlier work this paper cites.
F. Chen, R. Ji, J. Su, D. Cao, and Y. Gao, “Predicting microblog sentiments via weakly supervised multimodal deep learning,” IEEE Transactions on Multimedia , vol. 20, no. 4, pp. 997–1007, 2018
2018
Earlier work this paper cites.
A. Zadeh, P. P. Liang, J. Vanbriesen, S. Poria, E. Tong, E. Cambria, M. Chen, and L. P. Morency, “Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,” in ACL , 2018, pp. 2236–2246
2018
Earlier work this paper cites.
T. Baltrusaitis, A. Zadeh, Y. C. Lim, and L.-P. Morency, “Openface 2.0: Facial behavior analysis toolkit,” in 2018 13th IEEE international conference on automatic face & gesture recognition (FG 2018) . IEEE, 2018, pp. 59–66
2018
Earlier work this paper cites.
Z. Liu, Y. Shen, P. P. Liang, A. Zadeh, and L. P. Morency, “Efficient low-rank multimodal fusion with modality-specific factors,” in ACL , 2018, pp. 2247–2256
2018
Earlier work this paper cites.
S. Mai, H. Hu, and S. Xing, “Divide, conquer and combine: Hierarchical feature fusion network with local and global perspectives for multimodal affective computing,” in ACL , Jul. 2019, pp. 481–492
2019
Earlier work this paper cites.
P. P. Liang, Z. Liu, Y.-H. H. Tsai, Q. Zhao, R. Salakhutdinov, and L.-P. Morency, “Learning representations from imperfect time series data via tensor rank regularization,” in ACL , 2019, pp. 1569–1576
2019
Earlier work this paper cites.
Y.-H. H. Tsai, S. Bai, P. P. Liang, J. Z. Kolter, L.-P. Morency, and R. Salakhutdinov, “Multimodal transformer for unaligned multimodal language sequences,” in ACL , Jul. 2019, pp. 6558–6569
2019
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in NAACL-HLT , 2019
2019
Earlier work this paper cites.
S. H. Dumpala, I. Sheikh, R. Chakraborty, and S. K. Kopparapu, “Audio-visual fusion for sentiment classification using cross-modal autoencoder,” in 32nd conference on neural information processing systems (NIPS 2018) , 2019, pp. 1–4
2019
Earlier work this paper cites.
J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
R. Lotfian and C. Busso, “Curriculum learning for speech emotion recognition from crowdsourced labels,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 4, pp. 815–826, 2019
2019
Earlier work this paper cites.
2019
Cited alongside, same era.
Y. H. H. Tsai, P. P. Liang, A. Zadeh, L. P. Morency, and R. Salakhutdinov, “Learning factorized multimodal representations,” in ICLR , 2019
2019
Cited alongside, same era.
W. Rahman, M. Hasan, S. Lee, A. Zadeh, C. Mao, L.-P. Morency, and E. Hoque, “Integrating multimodal information in large pretrained transformers,” ACL , vol. 2020, pp. 2359–2369, 2020
2020
Cited alongside, same era.
Y.-H. H. Tsai, M. Q. Ma, M. Yang, R. Salakhutdinov, and L.-P. Morency, “Multimodal routing: Improving local and global interpretability of multimodal language analysis,” in Proceedings of the Conference on Empirical Methods in Natural Language Processing. Conference on Empirical Methods in Natural Language Processing , vol. 2020. NIH Public Access, 2020, p. 1823
2020
S. Shankar, “Multimodal fusion via cortical network inspired losses,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2022, pp. 1167–1178
2022
Later among the works it cites.
C.-Y. Chuang, R. D. Hjelm, X. Wang, V. Vineet, N. Joshi, A. Torralba, S. Jegelka, and Y. Song, “Robust contrastive learning against noisy views,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 16 670–16 681
2022
Later among the works it cites.
S. Goel, H. Bansal, S. Bhatia, R. Rossi, V. Vinay, and A. Grover, “Cyclip: Cyclic contrastive language-image pretraining,” Advances in Neural Information Processing Systems , vol. 35, pp. 6704–6719, 2022
2022
Later among the works it cites.
J. Zeng, J. Zhou, and T. Liu, “Mitigating inconsistencies in multimodal sentiment analysis under uncertain missing modalities,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , 2022, pp. 2924–2934
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Z. Sun, P. Sarma, W. Sethares, and Y. Liang, “Learning relationships between text, audio, and video via deep canonical correlation for multimodal language analysis,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 05, 2020, pp. 8992–8999
2020
Cited alongside, same era.
2020
Cited alongside, same era.
N. Pielawski, E. Wetzer, J. Öfverstedt, J. Lu, C. Wählby, J. Lindblad, and N. Sladoje, “Comir: Contrastive multimodal image representation for registration,” Advances in neural information processing systems , vol. 33, pp. 18 433–18 444, 2020
2020
Cited alongside, same era.
D. Hazarika, R. Zimmermann, and S. Poria, “Misa: Modality-invariant and -specific representations for multimodal sentiment analysis,” ACM MM , pp. 1122–1131, 2020
2020
Cited alongside, same era.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” The Journal of Machine Learning Research , vol. 21, no. 1, pp. 5485–5551, 2020
2020
Cited alongside, same era.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in International conference on machine learning . PMLR, 2020, pp. 1597–1607
2020
Cited alongside, same era.
W. Yu, H. Xu, F. Meng, Y. Zhu, Y. Ma, J. Wu, J. Zou, and K. Yang, “Ch-sims: A chinese multimodal sentiment analysis dataset with fine-grained annotation of modality,” in Proceedings of the 58th annual meeting of the association for computational linguistics , 2020, pp. 3718–3727
2020
Cited alongside, same era.
K. Yang, H. Xu, and K. Gao, “Cm-bert: Cross-modal bert for text-audio sentiment analysis,” in ACM MM , 2020, pp. 521–528
2020
Cited alongside, same era.
2022
Later among the works it cites.
G. Hu, T.-E. Lin, Y. Zhao, G. Lu, Y. Wu, and Y. Li, “Unimse: Towards unified multimodal sentiment analysis and emotion recognition,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , 2022, pp. 7837–7851
2022
Later among the works it cites.
R. Lin and H. Hu, “Multimodal contrastive learning via uni-modal coding and cross-modal prediction for multimodal sentiment analysis,” in Findings of the Association for Computational Linguistics: EMNLP 2022 , 2022, pp. 511–523
2022
Later among the works it cites.
H. Sun, J. Liu, Y.-W. Chen, and L. Lin, “Modality-invariant temporal representation learning for multimodal sentiment classification,” Information Fusion , vol. 91, pp. 504–514, 2023
2023
Later among the works it cites.
H. Sun, Y.-W. Chen, and L. Lin, “Tensorformer: A tensor-based multimodal transformer for multimodal sentiment analysis and depression detection,” IEEE Transactions on Affective Computing , vol. 14, no. 4, pp. 2776–2786, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Hwang and J.-H. Kim, “Self-supervised unimodal label generation strategy using recalibrated modality representations for multimodal sentiment analysis,” in Findings of the Association for Computational Linguistics: EACL 2023 , 2023, pp. 35–46
2023
Later among the works it cites.
Y. Tu, B. Zhang, Y. Li, L. Liu, J. Li, Y. Wang, C. Wang, and C. R. Zhao, “Learning from noisy labels with decoupled meta label purifier,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 19 934–19 943
2023
Later among the works it cites.
L. Zhu, Z. Zhu, C. Zhang, Y. Xu, and X. Kong, “Multimodal sentiment analysis based on fusion methods: A survey,” Information Fusion , vol. 95, pp. 306–325, 2023
2023
Later among the works it cites.
P. Koromilas, M. A. Nicolaou, T. Giannakopoulos, and Y. Panagakis, “Mmatr: A lightweight approach for multimodal sentiment analysis based on tensor methods,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
J. Zeng, J. Zhou, and C. Huang, “Exploring semantic relations for social media sentiment analysis,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2023
2023
Later among the works it cites.
K. Kim and S. Park, “Aobert: All-modalities-in-one bert for multimodal sentiment analysis,” Information Fusion , vol. 92, pp. 37–45, 2023
2023
Later among the works it cites.
Y. Sun, S. Mai, and H. Hu, “Learning to learn better unimodal representations via adaptive multimodal meta-learning,” IEEE Transactions on Affective Computing , vol. 14, no. 3, pp. 2209–2223, 2023
2023
Later among the works it cites.
S. Mai, Y. Sun, Y. Zeng, and H. Hu, “Excavating multimodal correlation for representation learning,” Information Fusion , vol. 91, pp. 542–555, 2023
2023
Later among the works it cites.
J. Yang, Y. Yu, D. Niu, W. Guo, and Y. Xu, “Confede: Contrastive feature decomposition for multimodal sentiment analysis,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2023, pp. 7617–7630
2023
Later among the works it cites.
S. Mai, Y. Zeng, S. Zheng, and H. Hu, “Hybrid contrastive learning of tri-modal representation for multimodal sentiment analysis,” IEEE Transactions on Affective Computing , vol. 14, no. 3, pp. 2276–2289, 2023
2023
Later among the works it cites.
S. Anand, N. K. Devulapally, S. D. Bhattacharjee, and J. Yuan, “Multi-label emotion analysis in conversation via multimodal knowledge distillation,” in Proceedings of the 31st ACM International Conference on Multimedia , 2023, pp. 6090–6100
2023
Later among the works it cites.
Q. Xia, F. Lee, and Q. Chen, “Tcc-net: A two-stage training method with contradictory loss and co-teaching based on meta-learning for learning with noisy labels,” Information Sciences , vol. 639, p. 119008, 2023
2023
Later among the works it cites.
X. Wang, J. Lyu, B.-G. Kim, B. Parameshachari, K. Li, and Q. Li, “Exploring multimodal multiscale features for sentiment analysis using fuzzy-deep neural network learning,” IEEE Transactions on Fuzzy Systems , pp. 1–15, 2024
2024
Closest in time.
2024
Closest in time.
S. Mai, Y. Sun, A. Xiong, Y. Zeng, and H. Hu, “Multimodal boosting: Addressing noisy modalities and identifying modality contribution,” IEEE Transactions on Multimedia , vol. 26, pp. 3018–3033, 2024
2024
Closest in time.