Fetching the paper…
Reading the bibliography…
Multimodal learning seeks to utilize data from multiple sources to improve the overall performance of downstream tasks.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Multimodal semi-supervised learning for image classification
Matthieu Guillaumin, Jakob Verbeek, and Cordelia Schmid · 2010
Earlier work this paper cites.
Indoor segmentation and support inference from RGBD images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus · 2012
Earlier work this paper cites.
Learning rich features from RGB-D images for object detection and segmentation
Saurabh Gupta, Ross Girshick, Pablo Arbeláez, and Jitendra Malik · 2014
Earlier work this paper cites.
Learning phrase representations using RNN encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Food-101–mining discriminative components with random forests
Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool · 2014
Earlier work this paper cites.
ModDrop: Adaptive multi-modal gesture recognition
Natalia Neverova, Christian Wolf, Graham Taylor, and Florian Nebout · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Recipe recognition with large multimodal food dataset
Xin Wang, Devinder Kumar, Nicolas Thome, Matthieu Cord, and Frederic Precioso · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, and Trevor Darrell · 2015
Earlier work this paper cites.
Layer normalization
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Ntu rgb+ d: A large scale dataset for 3d human activity analysis
Amir Shahroudy, Jun Liu, Tian-Tsong Ng, and Gang Wang · 2016
Earlier work this paper cites.
Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages
Amir Zadeh, Rowan Zellers, Eli Pincus, and Louis-Philippe Morency · 2016
Earlier work this paper cites.
FuseNet: Incorporating depth into semantic segmentation via fusion-based CNN architecture
Caner Hazirbas, Lingni Ma, Csaba Domokos, and Daniel Cremers · 2016
Earlier work this paper cites.
Multi-scale context aggregation by dilated convolutions
Fisher Yu and Vladlen Koltun · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
A survey of multimodal sentiment analysis
Mohammad Soleymani, David Garcia, Brendan Jou, Björn Schuller, Shih-Fu Chang, and Maja Pantic · 2017
Earlier work this paper cites.
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
Antti Tarvainen and Harri Valpola · 2017
Earlier work this paper cites.
MFNet: Towards real-time semantic segmentation for autonomous vehicles with multi-spectral scenes
Qishen Ha, Kohei Watanabe, Takumi Karasawa, Yoshitaka Ushiku, and Tatsuya Harada · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Tensor fusion network for multimodal sentiment analysis
Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency · 2017
Earlier work this paper cites.
Multimodal machine learning: A survey and taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency · 2018
Earlier work this paper cites.
3D cGAN based cross-modality MR image synthesis for brain tumor segmentation
Biting Yu, Luping Zhou, Lei Wang, Jurgen Fripp, and Pierrick Bourgeat · 2018
Earlier work this paper cites.
Group normalization
Yuxin Wu and Kaiming He · 2018
Earlier work this paper cites.
Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph
AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency · 2018
Earlier work this paper cites.
Efficient yet deep convolutional neural networks for semantic segmentation
Sharif Amit Kamran and Ali Shihab Sabbir · 2018
Earlier work this paper cites.
Memory fusion network for multi-view sequential learning
Amir Zadeh, Paul Pu Liang, Navonil Mazumder, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency · 2018
Earlier work this paper cites.
Efficient low-rank multimodal fusion with modality-specific factors
Zhun Liu, Ying Shen, Varun Bharadhwaj Lakshminarasimhan, Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency · 2018
Earlier work this paper cites.
Modality distillation with multiple stream networks for action recognition
Nuno C Garcia, Pietro Morerio, and Vittorio Murino · 2018
Earlier work this paper cites.
Graph distillation for action detection with privileged modalities
Zelun Luo, Jun-Ting Hsieh, Lu Jiang, Juan Carlos Niebles, and Li Fei-Fei · 2018
Earlier work this paper cites.
A multimodal fusion approach for image captioning
Dexin Zhao, Zhi Chang, and Shutao Guo · 2019
Cited alongside, same era.
Multimodal transformer with multi-view visual representation for image captioning
Jun Yu, Jing Li, Zhou Yu, and Qingming Huang · 2019
Cited alongside, same era.
Missing MRI pulse sequence synthesis using multi-modal generative adversarial network
Anmol Sharma and Ghassan Hamarneh · 2019
Cited alongside, same era.
Hetero-modal variational encoder-decoder for joint modality completion and segmentation
Reuben Dorent, Samuel Joutard, Marc Modat, Sébastien Ourselin, and Tom Vercauteren · 2019
Cited alongside, same era.
A unified representation network for segmentation with missing modalities
Kenneth Lau, Jonas Adler, and Jens Sjölund · 2019
Cited alongside, same era.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Later among the works it cites.
Robust fine-tuning of zero-shot models
Mitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li, Simon Kornblith, Rebecca Roelofs, Raphael Gontijo Lopes, Hannaneh Hajishirzi, Ali Farhadi, and Hongseok Namkoong · 2022
Later among the works it cites.
MM-align: Learning optimal transport-based alignment dynamics for fast and accurate inference on missing modality sequences
Wei Han, Hui Chen, Min-Yen Kan, and Soujanya Poria · 2022
Later among the works it cites.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin A Raffel · 2022
Later among the works it cites.
Towards a unified view of parameter-efficient transfer learning
Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, and Graham Neubig · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
EmbraceNet: A robust deep learning architecture for multimodal classification
Jun-Ho Choi and Jong-Seok Lee · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for nlp
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 2019
Cited alongside, same era.
Multimodal transformer for unaligned multimodal language sequences
Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
RTFNet: RGB-Thermal Fusion Network for Semantic Segmentation of Urban Scenes
Yuxiang Sun, Weixun Zuo, and Ming Liu · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova · 2019
Cited alongside, same era.
Bi-directional cross-modality feature propagation with separation-and-aggregation gate for RGB-D semantic segmentation
Xiaokang Chen, Kwan-Yee Lin, Jingbo Wang, Wayne Wu, Chen Qian, Hongsheng Li, and Gang Zeng · 2020
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Later among the works it cites.
Scaling & shifting your features: A new baseline for efficient model tuning
Dongze Lian, Daquan Zhou, Jiashi Feng, and Xinchao Wang · 2022
Later among the works it cites.
Krona: Parameter efficient tuning with kronecker adapter
Ali Edalati, Marzieh Tahaei, Ivan Kobyzev, Vahid Partovi Nia, James J Clark, and Mehdi Rezagholizadeh · 2022
Later among the works it cites.
BitFit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Elad Ben Zaken, Yoav Goldberg, and Shauli Ravfogel · 2022
Later among the works it cites.
Multimodal material segmentation
Yupeng Liang, Ryosuke Wakaki, Shohei Nobuhara, and Ko Nishino · 2022
Later among the works it cites.
Decoupling and recoupling spatiotemporal representation for rgb-d-based motion recognition
Benjia Zhou, Pichao Wang, Jun Wan, Yanyan Liang, Fan Wang, Du Zhang, Zhen Lei, Hao Li, and Rong Jin · 2022
Later among the works it cites.
Multimodal learning with transformers: A survey
Peng Xu, Xiatian Zhu, and David A Clifton · 2023
Closest in time.
Multimodal fusion transformer for remote sensing image classification
Swalpa Kumar Roy, Ankur Deria, Danfeng Hong, Behnood Rasti, Antonio Plaza, and Jocelyn Chanussot · 2023
Closest in time.
On robustness in multimodal learning
Brandon McKinzie, Joseph Cheng, Vaishaal Shankar, Yinfei Yang, Jonathon Shlens, and Alexander Toshev · 2023
Closest in time.
Towards good practices for missing modality robust action recognition
Sangmin Woo, Sumin Lee, Yeonju Park, Muhammad Adi Nugroho, and Changick Kim · 2023
Closest in time.
Multimodal prompting with missing modalities for visual recognition
Yi-Lun Lee, Yi-Hsuan Tsai, Wei-Chen Chiu, and Chen-Yu Lee · 2023
Closest in time.
What makes for robust multi-modal models in the face of missing modalities?
Siting Li, Chenzhuang Du, Yue Zhao, Yu Huang, and Hang Zhao · 2023
Closest in time.
SpiderMesh: Spatial-aware demand-guided recursive meshing for RGB-T semantic segmentation
Siqi Fan, Zhe Wang, Yan Wang, and Jingjing Liu · 2023
Closest in time.
Multi-modal learning with missing modality via shared-specific feature modelling
Hu Wang, Yuanhong Chen, Congbo Ma, Jodie Avery, Louise Hull, and Gustavo Carneiro · 2023
Closest in time.
Variational probabilistic fusion network for RGB-T semantic segmentation
Baihong Lin, Zengrong Lin, Yulan Guo, Yulan Zhang, Jianxiao Zou, and Shicai Fan · 2023
Closest in time.
Parameter-efficient fine-tuning of large-scale pre-trained language models
Ning Ding, Yujia Qin, Guang Yang, Fuchao Wei, Zonghan Yang, Yusheng Su, Shengding Hu, Yulin Chen, Chi-Min Chan, Weize Chen, et al · 2023
Closest in time.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer · 2023
Closest in time.
Parameter-efficient model adaptation for vision transformers
Xuehai He, Chunyuan Li, Pengchuan Zhang, Jianwei Yang, and Xin Eric Wang · 2023
Closest in time.
Delivering arbitrary-modal semantic segmentation
Jiaming Zhang, Ruiping Liu, Hao Shi, Kailun Yang, Simon Reiß, Kunyu Peng, Haodong Fu, Kaiwei Wang, and Rainer Stiefelhagen · 2023
Closest in time.
A unified multimodal de-and re-coupling framework for rgb-d motion recognition
Benjia Zhou, Pichao Wang, Jun Wan, Yanyan Liang, and Fan Wang · 2023
Closest in time.
Mitigating modality discrepancies for RGB-T semantic segmentation
Shenlu Zhao, Yichen Liu, Qiang Jiao, Qiang Zhang, and Jungong Han · 2023
Closest in time.
Mmsformer: Multimodal transformer for material and semantic segmentation
Md Kaykobad Reza, Ashley Prater-Bennette, and M. Salman Asif · 2024
Closest in time.
Complementary random masking for rgb-thermal semantic segmentation
Ukcheol Shin, Kyunghyun Lee, In So Kweon, and Jean Oh · 2024
Closest in time.
Missing modality robustness in semi-supervised multi-modal semantic segmentation
Harsh Maheshwari, Yen-Cheng Liu, and Zsolt Kira · 2024
Closest in time.
Exploring missing modality in multimodal egocentric datasets
Merey Ramazanova, Alejandro Pardo, Humam Alwassel, and Bernard Ghanem · 2024
Closest in time.