Fetching the paper…
Reading the bibliography…
With the growing success of multi-modal learning, research on the robustness of multi-modal models, especially when facing situations with missing modalities, is receiving increased attention.
On single source robustness in deep fusion models
Taewan Kim and Joydeep Ghosh · 1906
Earlier work this paper cites.
Yonglong Tian, Dilip Krishnan, and Phillip Isola · 1906
Earlier work this paper cites.
CLEVRER: collision events for video representation and reasoning
Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, and Joshua B. Tenenbaum · 1910
Earlier work this paper cites.
Listen to look: Action recognition by previewing audio
Ruohan Gao, Tae-Hyun Oh, Kristen Grauman, and Lorenzo Torresani · 1912
Earlier work this paper cites.
Bayes error estimation using parzen and k-nn procedures
Keinosuke Fukunaga and Donald M Hummels · 1987
Earlier work this paper cites.
Relations between entropy and error probability
M. Feder and N. Merhav · 1994
Earlier work this paper cites.
Elements of information theory
Thomas M Cover · 1999
Earlier work this paper cites.
A co-regularization approach to semi-supervised learning with multiple views
Vikas Sindhwani, Partha Niyogi, and Mikhail Belkin · 2005
Earlier work this paper cites.
Investigating vulnerability to adversarial examples on multimodal data fusion in deep learning
Youngjoon Yu, Hong Joo Lee, Byeong Cheon Kim, Jung Uk Kim, and Yong Man Ro · 2005
Earlier work this paper cites.
Effects of missing data in social networks
Gueorgi Kossinets · 2006
Earlier work this paper cites.
Demystifying self-supervised learning: An information-theoretical framework
Yao-Hung Hubert Tsai, Yue Wu, Ruslan Salakhutdinov, and Louis-Philippe Morency · 2006
Earlier work this paper cites.
Multi-view regression via canonical correlation analysis
Sham M Kakade and Dean P Foster · 2007
Earlier work this paper cites.
Multi-domain image completion for random missing input data
Liyue Shen, Wentao Zhu, Xiaosong Wang, Lei Xing, John M. Pauly, Baris Turkbey, Stephanie Anne Harmon, Thomas Hogue Sanford, Sherif Mehralivand, Peter L. Choyke, Bradford J. Wood, and Daguang Xu · 2007
Earlier work this paper cites.
An information theoretic framework for multi-view learning
Karthik Sridharan and Sham M. Kakade · 2008
Earlier work this paper cites.
Learning from multiple partially observed views - an application to multilingual text categorization
Massih R. Amini, Nicolas Usunier, and Cyril Goutte · 2009
Earlier work this paper cites.
Functional properties of minimum mean-square error and mutual information
Yihong Wu and Sergio Verdu · 2011
Earlier work this paper cites.
A closer look at the robustness of vision-and-language pre-trained models
Linjie Li, Zhe Gan, and Jingjing Liu · 2012
Earlier work this paper cites.
Multi-source learning for joint analysis of incomplete multi-modality neuroimaging data
Lei Yuan, Yalin Wang, Paul M. Thompson, Vaibhav A. Narayan, and Jieping Ye · 2012
Earlier work this paper cites.
A survey on multi-view learning
Chang Xu, Dacheng Tao, and Chao Xu · 2013
Earlier work this paper cites.
VQA: visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Multimodal deep learning for robust RGB-D object recognition
Andreas Eitel, Jost Tobias Springenberg, Luciano Spinello, Martin A. Riedmiller, and Wolfram Burgard · 2015
Earlier work this paper cites.
Information theory: a tutorial introduction
James V Stone · 2015
Earlier work this paper cites.
Recipe recognition with large multimodal food dataset
Xin Wang, Devinder Kumar, Nicolas Thome, Matthieu Cord, and Frederic Precioso · 2015
Earlier work this paper cites.
Convolutional two-stream network fusion for video action recognition
Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross B. Girshick · 2016
Cited alongside, same era.
Learning common and specific features for rgb-d semantic segmentation with deconvolutional networks
Jinghua Wang, Zhenhua Wang, Dacheng Tao, Simon See, and Gang Wang · 2016
Cited alongside, same era.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Cited alongside, same era.
Gated multimodal units for information fusion
John Arevalo, Thamar Solorio, Manuel Montes-y Gómez, and Fabio A González · 2017
Cited alongside, same era.
Modality dropout for improved performance-driven talking faces
Ahmed Hussen Abdelaziz, Barry-John Theobald, Paul Dixon, Reinhard Knothe, Nicholas Apostoloff, and Sachin Kajareker · 2020
Later among the works it cites.
Tcgm: An information-theoretic framework for semi-supervised multi-modality learning
Xinwei Sun, Yilun Xu, Peng Cao, Yuqing Kong, Lingjing Hu, Shanghang Zhang, and Yizhou Wang · 2020
Later among the works it cites.
What makes training multi-modal classification networks hard?
Weiyao Wang, Du Tran, and Matt Feiszli · 2020
Later among the works it cites.
Cooperative learning for multi-view analysis, 2021
Daisy Yi Ding, Balasubramanian Narasimhan, and Robert Tibshirani · 2021
Later among the works it cites.
Trusted multi-view classification
Zongbo Han, Changqing Zhang, Huazhu Fu, and Joey Tianyi Zhou · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tadas Baltrusaitis, Chaitanya Ahuja, and Louis-Philippe Morency · 2017
Cited alongside, same era.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Cited alongside, same era.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Cited alongside, same era.
Look, imagine and match: Improving textual-visual cross-modal retrieval with generative models
Jiuxiang Gu, Jianfei Cai, Shafiq R. Joty, Li Niu, and Gang Wang · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Vigan: Missing view imputation with generative adversarial networks
Chao Shang, Aaron Palmer, Jiangwen Sun, Ko-Shin Chen, Jin Lu, and Jinbo Bi · 2017
Cited alongside, same era.
Missing modalities imputation via cascaded residual autoencoder
Luan Tran, Xiaoming Liu, Jiayu Zhou, and Rong Jin · 2017
Cited alongside, same era.
What makes multimodal learning better than single (provably)
Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen, Hang Zhao, and Longbo Huang · 2021
Later among the works it cites.
Vilt: Vision-and-language transformer without convolution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim · 2021
Later among the works it cites.
Multibench: Multiscale benchmarks for multimodal representation learning
Paul Pu Liang, Yiwei Lyu, Xiang Fan, Zetian Wu, Yun Cheng, Jason Wu, Leslie Yufan Chen, Peter Wu, Michelle A Lee, Yuke Zhu, et al · 2021
Later among the works it cites.
COMPLETER: incomplete multi-view clustering via contrastive prediction
Yijie Lin, Yuanbiao Gou, Zitao Liu, Boyun Li, Jiancheng Lv, and Xi Peng · 2021
Later among the works it cites.
Smil: Multimodal learning with severely missing modality
Mengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov, Cathy Wu, and Xi Peng · 2021
Later among the works it cites.
Are VQA systems rad? measuring robustness to augmented data with focused interventions
Daniel Rosenberg, Itai Gat, Amir Feder, and Roi Reichart · 2021
Later among the works it cites.
Can audio-visual integration strengthen robustness under multimodal attacks?
Yapeng Tian and Chenliang Xu · 2021
Later among the works it cites.
Contrastive learning, multi-view redundancy, and linear models
Christopher Tosh, Akshay Krishnamurthy, and Daniel Hsu · 2021
Later among the works it cites.
On robustness to missing video for audiovisual speech recognition
Oscar Chang, Otavio de Pinho Forin Braga, Hank Liao, Dmitriy Dima Serdyuk, and Olivier Siohan · 2022
Later among the works it cites.
Yu Huang, Junyang Lin, Chang Zhou, Hongxia Yang, and Longbo Huang · 2022
Later among the works it cites.
Dual contrastive prediction for incomplete multi-view representation learning
Yijie Lin, Yuanbiao Gou, Xiaotian Liu, Jinfeng Bai, Jiancheng Lv, and Xi Peng · 2022
Later among the works it cites.
Are multimodal transformers robust to missing modality?, 2022
Mengmeng Ma, Jian Ren, Long Zhao, Davide Testuggine, and Xi Peng · 2022
Later among the works it cites.
Balanced multimodal learning via on-the-fly gradient modulation
Xiaokang Peng, Yake Wei, Andong Deng, Dong Wang, and Di Hu · 2022
Later among the works it cites.
Reproducible scaling laws for contrastive language-image learning
Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuhmann, Ludwig Schmidt, and Jenia Jitsev · 2023
Closest in time.
On uni-modal feature learning in supervised multi-modal learning
Chenzhuang Du, Jiaye Teng, Tingle Li, Yichen Liu, Tianyuan Yuan, Yue Wang, Yang Yuan, and Hang Zhao · 2023
Closest in time.
Imagebind: One embedding space to bind them all
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra · 2023
Closest in time.
Multi-modal learning with missing modality via shared-specific feature modelling
Hu Wang, Yuanhong Chen, Congbo Ma, Jodie Avery, Louise Hull, and Gustavo Carneiro · 2023
Closest in time.
Meta-transformer: A unified framework for multimodal learning
Yiyuan Zhang, Kaixiong Gong, Kaipeng Zhang, Hongsheng Li, Yu Qiao, Wanli Ouyang, and Xiangyu Yue · 2023
Closest in time.