Fetching the paper…
Reading the bibliography…
Multimodal representation learning aims to capture both shared and complementary semantic information across multiple modalities.
Pattern Recognition and Machine Learning
Christopher M Bishop · 2006
Earlier work this paper cites.
Iemocap: Interactive emotional dyadic motion capture database
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan · 2008
Earlier work this paper cites.
Multi-marginal optimal transport: theory and applications
Brendan Pass · 2015
Earlier work this paper cites.
Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages
Amir Zadeh, Rowan Zellers, Eli Pincus, and Louis-Philippe Morency · 2016
Earlier work this paper cites.
Learning factorized multimodal representations
Yao-Hung Hubert Tsai, Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2018
Earlier work this paper cites.
Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph
AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency · 2018
Earlier work this paper cites.
Computational optimal transport: With applications to data science
Gabriel Peyré and Marco Cuturi · 2019
Earlier work this paper cites.
wav2vec: Unsupervised pre-training for speech recognition
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli · 2019
Earlier work this paper cites.
Multimodal transformer for unaligned multimodal language sequences
Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
Misa: Modality-invariant and-specific representations for multimodal sentiment analysis
Devamanyu Hazarika, Roger Zimmermann, and Soujanya Poria · 2020
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2020
Earlier work this paper cites.
Ch-sims: A chinese multimodal sentiment analysis dataset with fine-grained annotation of modality
Wenmeng Yu, Hua Xu, Fanyang Meng, Yilin Zhu, Yixiao Ma, Jiele Wu, Jiyun Zou, and Kaicheng Yang · 2020
Earlier work this paper cites.
Attention is not enough: Mitigating the distribution discrepancy in asynchronous multimodal sequence fusion
Tao Liang, Guosheng Lin, Lei Feng, Yan Zhang, and Fengmao Lv · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Learning modality-specific representations with self-supervised multi-task learning for multimodal sentiment analysis
Wenmeng Yu, Hua Xu, Ziqi Yuan, and Jiele Wu · 2021
Earlier work this paper cites.
Learning deep global multi-scale and local attention features for facial expression recognition in the wild
Zengqun Zhao, Qingshan Liu, and Shanmin Wang · 2021
Earlier work this paper cites.
Multimodal dynamics: Dynamical fusion for trustworthy multimodal classification
Zongbo Han, Fan Yang, Junzhou Huang, Changqing Zhang, and Jianhua Yao · 2022
Earlier work this paper cites.
Modeling intra-and inter-modal relations: Hierarchical graph contrastive learning for multimodal sentiment analysis
Zijie Lin, Bin Liang, Yunfei Long, Yixue Dang, Min Yang, Min Zhang, and Ruifeng Xu · 2022
Earlier work this paper cites.
Disentangled multimodal representation learning for recommendation
Fan Liu, Huilin Chen, Zhiyong Cheng, Anan Liu, Liqiang Nie, and Mohan Kankanhalli · 2022
Cited alongside, same era.
Cubemlp: An mlp-based model for multimodal sentiment analysis and depression estimation
Hao Sun, Hongyi Wang, Jiaqing Liu, Yen-Wei Chen, and Lanfen Lin · 2022
Cited alongside, same era.
Multimodal disentanglement variational autoencoders for zero-shot cross-modal retrieval
Jialin Tian, Kai Wang, Xing Xu, Zuo Cao, Fumin Shen, and Heng Tao Shen · 2022
Cited alongside, same era.
Cross-modal enhancement network for multimodal sentiment analysis
Di Wang, Shuai Liu, Quan Wang, Yumin Tian, Lihuo He, and Xinbo Gao · 2022
Cited alongside, same era.
Disentangled representation learning for multimodal emotion recognition
Dingkang Yang, Shuai Huang, Haopeng Kuang, Yangtao Du, and Lihua Zhang · 2022
Cited alongside, same era.
Clip-adapter: Better vision-language models with feature adapters
Peng Gao, Shijie Geng, Renrui Zhang, Teli Ma, Rongyao Fang, Yongfeng Zhang, Hongsheng Li, and Yu Qiao · 2024
Later among the works it cites.
A novel decoupled prototype completion network for incomplete multimodal emotion recognition
Zhangfeng Hu, Wenming Zheng, Yuan Zong, Mengting Wei, Xingxun Jiang, and Mengxin Shi · 2024
Later among the works it cites.
Reconboost: Boosting can achieve modality reconcilement
Cong Hua, Qianqian Xu, Shilong Bao, Zhiyong Yang, and Qingming Huang · 2024
Later among the works it cites.
Emma: End-to-end multimodal model for autonomous driving
Jyh-Jing Hwang, Runsheng Xu, Hubert Lin, Wei-Chih Hung, Jingwei Ji, Kristy Choi, Di Huang, Tong He, Paul Covington, Benjamin Sapp, et al · 2024
Later among the works it cites.
On-the-fly modulation for balanced multimodal learning
Yake Wei, Di Hu, Henghui Du, and Ji-Rong Wen · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ringki Das and Thoudam Doren Singh · 2023
Cited alongside, same era.
Pmr: Prototypical modal rebalance for multimodal learning
Yunfeng Fan, Wenchao Xu, Haozhao Wang, Junxiao Wang, and Song Guo · 2023
Cited alongside, same era.
Aobert: All-modalities-in-one bert for multimodal sentiment analysis
Kyeonghun Kim and Sanghyun Park · 2023
Cited alongside, same era.
Decoupled multimodal distilling for emotion recognition
Yong Li, Yuanzhi Wang, and Zhen Cui · 2023
Cited alongside, same era.
Gcnet: Graph completion network for incomplete multimodal learning in conversation
Zheng Lian, Lan Chen, Licai Sun, Bin Liu, and Jianhua Tao · 2023
Cited alongside, same era.
Learning discriminative multi-relation representations for multimodal sentiment analysis
Zemin Tang, Qi Xiao, Xu Zhou, Yangfan Li, Cen Chen, and Kenli Li · 2023
Cited alongside, same era.
Distribution-consistent modal recovering for incomplete multimodal learning
Yuanzhi Wang, Zhen Cui, and Yong Li · 2023
Cited alongside, same era.
Later among the works it cites.
Achieving cross modal generalization with multimodal unified representation
Yan Xia, Hai Huang, Jieming Zhu, and Zhou Zhao · 2024
Later among the works it cites.
Towards multimodal sentiment analysis debiasing via bias purification
Dingkang Yang, Mingcheng Li, Dongling Xiao, Yang Liu, Kun Yang, Zhaoyu Chen, Yuzheng Wang, Peng Zhai, Ke Li, and Lihua Zhang · 2024
Later among the works it cites.
Disentanglement translation network for multimodal sentiment analysis
Ying Zeng, Wenjun Yan, Sijie Mai, and Haifeng Hu · 2024
Later among the works it cites.
A multi-level alignment and cross-modal unified semantic graph refinement network for conversational emotion recognition
Xiaoheng Zhang, Weigang Cui, Bin Hu, and Yang Li · 2024
Later among the works it cites.
Unraveling cross-modality knowledge conflicts in large vision-language models
Tinghui Zhu, Qin Liu, Fei Wang, Zhengzhong Tu, and Muhao Chen · 2024
Later among the works it cites.
Classifier-guided gradient modulation for enhanced multimodal learning
Zirun Guo, Tao Jin, Jingyuan Chen, and Zhou Zhao · 2025
Closest in time.
Dolphins: Multimodal language model for driving
Yingzi Ma, Yulong Cao, Jiachen Sun, Marco Pavone, and Chaowei Xiao · 2025
Closest in time.
Dyncim: Dynamic curriculum for imbalanced multimodal learning
Chengxuan Qian, Kai Han, Jingchao Wang, Zhenlong Yuan, Rui Qian, Chongwen Lyu, Jun Chen, and Zhe Liu · 2025
Closest in time.
Dlf: Disentangled-language-focused multimodal sentiment analysis
Pan Wang, Qiang Zhou, Yawen Wu, Tianlong Chen, and Jingtong Hu · 2025
Closest in time.
Diagnosing and re-learning for balanced multimodal learning
Yake Wei, Siwei Li, Ruoxuan Feng, and Di Hu · 2025
Closest in time.
Re-align: Aligning vision language models via retrieval-augmented direct preference optimization
Shuo Xing, Yuping Wang, Peiran Li, Ruizheng Bai, Yueqi Wang, Chengxuan Qian, Huaxiu Yao, and Zhengzhong Tu · 2025
Closest in time.
Triple disentangled representation learning for multimodal affective analysis
Ying Zhou, Xuefeng Liang, Han Chen, Yin Zhao, Xin Chen, and Lida Yu · 2025
Closest in time.