Fetching the paper…
Reading the bibliography…
Multimodal learning integrates complementary information from diverse modalities to enhance the decision-making process.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K Soomro · 2012
Earlier work this paper cites.
Crema-d: Crowd-sourced emotional multimodal actors dataset
Houwei Cao, David G Cooper, Michael K Keutmann, Ruben C Gur, Ani Nenkova, and Ragini Verma · 2014
Earlier work this paper cites.
Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages
Amir Zadeh, Rowan Zellers, Eli Pincus, and Louis-Philippe Morency · 2016
Earlier work this paper cites.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Earlier work this paper cites.
Efficient large-scale multi-modal classification
Douwe Kiela, Edouard Grave, Armand Joulin, and Tomas Mikolov · 2018
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville · 2018
Earlier work this paper cites.
Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph
AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency · 2018
Earlier work this paper cites.
What makes training multi-modal classification networks hard?
Weiyao Wang, Du Tran, and Matt Feiszli · 2020
Earlier work this paper cites.
Ch-sims: A chinese multimodal sentiment analysis dataset with fine-grained annotation of modality
Wenmeng Yu, Hua Xu, Fanyang Meng, Yilin Zhu, Yixiao Ma, Jiele Wu, Jiyun Zou, and Kaicheng Yang · 2020
Earlier work this paper cites.
Improving multi-modal learning with uni-modal teachers
Chenzhuang Du, Tingle Li, Yichen Liu, Zixin Wen, Tianyu Hua, Yue Wang, and Hang Zhao · 2021
Earlier work this paper cites.
Competence-based multimodal curriculum learning for medical report generation
Fenglin Liu, Shen Ge, and Xian Wu · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
A survey on curriculum learning
Xin Wang, Yudong Chen, and Wenwu Zhu · 2021
Cited alongside, same era.
Learning modality-specific representations with self-supervised multi-task learning for multimodal sentiment analysis
Wenmeng Yu, Hua Xu, Ziqi Yuan, and Jiele Wu · 2021
Cited alongside, same era.
Characterizing and overcoming the greedy nature of learning in multi-modal deep neural networks
Nan Wu, Stanislaw Jastrzebski, Kyunghyun Cho, and Krzysztof J Geras · 2022
Cited alongside, same era.
Predictive dynamic fusion
Bing Cao, Yinan Xia, Yi Ding, Changqing Zhang, and Qinghua Hu · 2024
Later among the works it cites.
Classifier-guided gradient modulation for enhanced multimodal learning
Zirun Guo, Tao Jin, Jingyuan Chen, and Zhou Zhao · 2024
Later among the works it cites.
Reconboost: Boosting can achieve modality reconcilement
Cong Hua, Qianqian Xu, Shilong Bao, Zhiyong Yang, and Qingming Huang · 2024
Later among the works it cites.
Emma: End-to-end multimodal model for autonomous driving
Jyh-Jing Hwang, Runsheng Xu, Hubert Lin, Wei-Chih Hung, Jingwei Ji, Kristy Choi, Di Huang, Tong He, Paul Covington, Benjamin Sapp, et al · 2024
Later among the works it cites.
Foundations & trends in multimodal machine learning: Principles, challenges, and open questions
Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Arindam Das, Sudip Das, Ganesh Sistu, Jonathan Horgan, Ujjwal Bhattacharya, Edward Jones, Martin Glavin, and Ciarán Eising · 2023
Cited alongside, same era.
Multimodal sentiment analysis: a survey of methods, trends, and challenges
Ringki Das and Thoudam Doren Singh · 2023
Cited alongside, same era.
Pmr: Prototypical modal rebalance for multimodal learning
Yunfeng Fan, Wenchao Xu, Haozhao Wang, Junxiao Wang, and Song Guo · 2023
Cited alongside, same era.
A review on methods and applications in multimodal deep learning
Summaira Jabeen, Xi Li, Muhammad Shoib Amin, Omar Bourahla, Songyuan Li, and Abdul Jabbar · 2023
Cited alongside, same era.
Towards balanced active learning for multimodal classification
Meng Shen, Yizheng Huang, Jianxiong Yin, Heqing Zou, Deepu Rajan, and Simon See · 2023
Cited alongside, same era.
Multimodal learning with transformers: A survey
Peng Xu, Xiatian Zhu, and David A Clifton · 2023
Cited alongside, same era.
Intra- and inter-modal curriculum for multimodal learning
Yuwei Zhou, Xin Wang, Hong Chen, Xuguang Duan, and Wenwu Zhu · 2023
Cited alongside, same era.
Deep imbalanced learning for multimodal emotion recognition in conversations
Tao Meng, Yuntao Shou, Wei Ai, Nan Yin, and Keqin Li · 2024
Later among the works it cites.
Enhancing visual-language modality alignment in large vision language models via self-improvement
Xiyao Wang, Jiuhai Chen, Zhaoyang Wang, Yuhang Zhou, Yiyang Zhou, Huaxiu Yao, Tianyi Zhou, Tom Goldstein, Parminder Bhatia, Furong Huang, et al · 2024
Later among the works it cites.
Mmpareto: boosting multimodal learning with innocent unimodal assistance
Yake Wei and Di Hu · 2024
Later among the works it cites.
Decalign: Hierarchical cross-modal alignment for decoupled multimodal representation learning
Chengxuan Qian, Shuo Xing, Shawn Li, Yue Zhao, and Zhengzhong Tu · 2025
Closest in time.
Re-align: Aligning vision language models via retrieval-augmented direct preference optimization
Shuo Xing, Yuping Wang, Peiran Li, Ruizheng Bai, Yueqi Wang, Chengxuan Qian, Huaxiu Yao, and Zhengzhong Tu · 2025
Closest in time.
Pure vision language action (vla) models: A comprehensive survey
Dapeng Zhang, Jin Sun, Chenghui Hu, Xiaoyan Wu, Zhenlong Yuan, Rui Zhou, Fei Shen, and Qingguo Zhou · 2025
Closest in time.