Fetching the paper…
Reading the bibliography…
Multimodal machine learning has gained significant attention in recent years due to its potential for integrating information from multiple modalities to enhance learning and decision-making processes.
Improving neural networks by preventing co-adaptation of feature detectors
Geoffrey E. Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2012
Earlier work this paper cites.
Crema-d: Crowd-sourced emotional multimodal actors dataset
Houwei Cao, David G. Cooper, Michael K. Keutmann, Ruben C. Gur, Ani Nenkova, and Ragini Verma · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks, 2015
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Actions transformations, 2016
Xiaolong Wang, Ali Farhadi, and Abhinav Gupta · 2016
Earlier work this paper cites.
Learning spatio-temporal representation with pseudo-3d residual networks, 2017
Zhaofan Qiu, Ting Yao, and Tao Mei · 2017
Cited alongside, same era.
Not just a black box: Learning important features through propagating activation differences, 2017
Avanti Shrikumar, Peyton Greenside, Anna Shcherbina, and Anshul Kundaje · 2017
Cited alongside, same era.
Towards better understanding of gradient-based attribution methods for deep neural networks, 2018
Marco Ancona, Enea Ceolini, Cengiz Öztireli, and Markus Gross · 2018
Cited alongside, same era.
Vggsound: A large-scale audio-visual dataset, 2020
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman · 2020
Cited alongside, same era.
What makes training multi-modal classification networks hard?, 2020
Weiyao Wang, Du Tran, and Matt Feiszli · 2020
Cited alongside, same era.
Greedy gradient ensemble for robust visual question answering, 2021
Xinzhe Han, Shuhui Wang, Chi Su, Qingming Huang, and Qi Tian · 2021
Later among the works it cites.
Improving multi-modal learning with uni-modal teachers, 2021
Chenzhuang Du, Tingle Li, Yichen Liu, Zixin Wen, Tianyu Hua, Yue Wang, and Hang Zhao · 2021
Later among the works it cites.
Balanced multimodal learning via on-the-fly gradient modulation, 2022
Xiaokang Peng, Yake Wei, Andong Deng, Dong Wang, and Di Hu · 2022
Later among the works it cites.
Maea: Multimodal attribution for embodied ai
Vidhi Jain, Jayant Sravan Tamarapalli, Sahiti Yerramilli, and Yonatan Bisk · 2023
Later among the works it cites.
Semantic augmentation in images using language
Sahiti Yerramilli, Jayant Sravan Tamarapalli, Tanmay Girish Kulkarni, Jonathan Francis, and Eric Nyberg · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…