Fetching the paper…
Reading the bibliography…
Research on multi-modal learning dominantly aligns the modalities in a unified space at training, and only a single one is taken for prediction at inference.
Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories
Li Fei-Fei, Rob Fergus, and Pietro Perona · 2004
Earlier work this paper cites.
Interactive multimodal robot programming
Soshi Iba, Christiaan JJ Paredis, and Pradeep K Khosla · 2005
Earlier work this paper cites.
Interactive multimodal learning environments: Special issue on interactive learning environments: Contemporary issues and trends
Roxana Moreno and Richard Mayer · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Multimodal deep learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng · 2011
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2014
Earlier work this paper cites.
A dataset and taxonomy for urban sound research
Justin Salamon, Christopher Jacoby, and Juan Pablo Bello · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Converting static image datasets to spiking neuromorphic datasets using saccades
Garrick Orchard, Ajinkya Jayawant, Gregory K Cohen, and Nitish Thakor · 2015
Earlier work this paper cites.
Esc: Dataset for environmental sound classification
Karol J Piczak · 2015
Earlier work this paper cites.
3d shapenets: A deep representation for volumetric shapes
Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Linguang Zhang, Xiaoou Tang, and Jianxiong Xiao · 2015
Earlier work this paper cites.
Cross modal distillation for supervision transfer
Saurabh Gupta, Judy Hoffman, and Jitendra Malik · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
Vidloc: A deep spatio-temporal model for 6-dof video-clip relocalization
Ronald Clark, Sen Wang, Andrew Markham, Niki Trigoni, and Hongkai Wen · 2017
Earlier work this paper cites.
Graph-based object classification for neuromorphic vision sensing
Yin Bi, Aaron Chadha, Alhabib Abbas, Eirina Bourtsoulatze, and Yiannis Andreopoulos · 2019
Earlier work this paper cites.
End-to-end learning of representations for asynchronous event-based data
Daniel Gehrig, Antonio Loquercio, Konstantinos G Derpanis, and Davide Scaramuzza · 2019
Earlier work this paper cites.
A comprehensive overhaul of feature distillation
Byeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park, Nojun Kwak, and Jin Young Choi · 2019
Earlier work this paper cites.
Audiocaps: Generating captions for audios in the wild
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim · 2019
Earlier work this paper cites.
Relational knowledge distillation
Wonpyo Park, Dongju Kim, Yan Lu, and Minsu Cho · 2019
Earlier work this paper cites.
An efficient approach to informative feature extraction from multimodal data
Lichen Wang, Jiaxiang Wu, Shao-Lun Huang, Lizhong Zheng, Xiangxiang Xu, Lin Zhang, and Junzhou Huang · 2019
Earlier work this paper cites.
Towards cross-modality medical image segmentation with online mutual knowledge distillation
Kang Li, Lequan Yu, Shujun Wang, and Pheng-Ann Heng · 2020
Earlier work this paper cites.
Knowledge distillation meets self-supervision
Guodong Xu, Ziwei Liu, Xiaoxiao Li, and Chen Change Loy · 2020
Earlier work this paper cites.
Progress and prospects of multimodal fusion methods in physical human–robot interaction: A review
Teng Xue, Weiming Wang, Jin Ma, Wenhai Liu, Zhenyu Pan, and Mingshuo Han · 2020
Earlier work this paper cites.
Knowledge transfer via dense cross-layer mutual-distillation
Anbang Yao and Dawei Sun · 2020
Earlier work this paper cites.
Knowledge as priors: Cross-modal knowledge generalization for datasets without superior knowledge
Long Zhao, Xi Peng, Yuxiao Chen, Mubbasir Kapadia, and Dimitris N Metaxas · 2020
Earlier work this paper cites.
Unified vision-language pre-training for image captioning and vqa
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason Corso, and Jianfeng Gao · 2020
Earlier work this paper cites.
Vlmo: Unified vision-language pre-training with mixture-of-modality-experts
Hangbo Bao, Wenhui Wang, Li Dong, Qiang Liu, Owais Khan Mohammed, Kriti Aggarwal, Subhojit Som, and Furu Wei · 2021
Earlier work this paper cites.
Clip2video: Mastering video-text retrieval via image clip
Han Fang, Pengfei Xiong, Luhui Xu, and Yu Chen · 2021
Cited alongside, same era.
Materials, actuators, and sensors for soft bioinspired robots
Mahdi Ilami, Hosain Bagheri, Reza Ahmed, E Olga Skowronek, and Hamid Marvi · 2021
Cited alongside, same era.
Llvip: A visible-infrared paired dataset for low-light vision
Xinyu Jia, Chuang Zhu, Minzhen Li, Wenqi Tang, and Wenli Zhou · 2021
Cited alongside, same era.
N-imagenet: Towards robust, fine-grained object recognition with event cameras
Junho Kim, Jaehyeok Bae, Gangin Park, Dongsu Zhang, and Young Min Kim · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Imagebind: One embedding space to bind them all
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra · 2023
Later among the works it cites.
Ziyu Guo, Renrui Zhang, Xiangyang Zhu, Yiwen Tang, Xianzheng Ma, Jiaming Han, Kexin Chen, Peng Gao, Xianzhi Li, Hongsheng Li, et al · 2023
Later among the works it cites.
Imagebind-llm: Multi-modality instruction tuning
Jiaming Han, Renrui Zhang, Wenqi Shao, Peng Gao, Peng Xu, Han Xiao, Kaipeng Zhang, Chris Liu, Song Wen, Ziyu Guo, et al · 2023
Later among the works it cites.
Vit-lens: Towards omni-modal representations
Weixian Lei, Yixiao Ge, Jianfeng Zhang, Dylan Sun, Kun Yi, Ying Shan, and Mike Zheng Shou · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Knowledge distillation and student-teacher learning for visual intelligence: A review and new outlooks
Lin Wang and Kuk-Jin Yoon · 2021
Cited alongside, same era.
Vinvl: Revisiting visual representations in vision-language models
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao · 2021
Cited alongside, same era.
Vlmo: Unified vision-language pre-training with mixture-of-modality-experts
Hangbo Bao, Wenhui Wang, Li Dong, Qiang Liu, Owais Khan Mohammed, Kriti Aggarwal, Subhojit Som, Songhao Piao, and Furu Wei · 2022
Cited alongside, same era.
Multimodal sensors and ml-based data fusion for advanced robots
Shengshun Duan, Qiongfeng Shi, and Jun Wu · 2022
Cited alongside, same era.
Audioclip: Extending clip to image, text and audio
Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel · 2022
Cited alongside, same era.
Clip2point: Transfer clip to point cloud classification with image-depth pre-training
Tianyu Huang, Bowen Dong, Yunhan Yang, Xiaoshui Huang, Rynson WH Lau, Wanli Ouyang, and Wangmeng Zuo · 2022
Cited alongside, same era.
Self-supervised visuo-tactile pretraining to locate and follow garment features
Justin Kerr, Huang Huang, Albert Wilcox, Ryan Hoque, Jeffrey Ichnowski, Roberto Calandra, and Ken Goldberg · 2022
Cited alongside, same era.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Later among the works it cites.
Ave-clip: Audioclip-based multi-window temporal transformer for audio visual event localization
Tanvir Mahmud and Diana Marculescu · 2023
Later among the works it cites.
Open vocabulary semantic segmentation with patch aligned contrastive learning
Jishnu Mukhoti, Tsung-Yu Lin, Omid Poursaeed, Rui Wang, Ashish Shah, Philip HS Torr, and Ser-Nam Lim · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Recent advancements in multimodal human–robot interaction
Hang Su, Wen Qi, Jiahao Chen, Chenguang Yang, Juan Sandoval, and Med Amine Laribi · 2023
Later among the works it cites.
iclip: Bridging image classification and contrastive language-image pre-training for visual recognition
Yixuan Wei, Yue Cao, Zheng Zhang, Houwen Peng, Zhuliang Yao, Zhenda Xie, Han Hu, and Baining Guo · 2023
Later among the works it cites.
Next-gpt: Any-to-any multimodal llm
Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua · 2023
Later among the works it cites.
Eventclip: Adapting clip for event-based object recognition
Ziyi Wu, Xudong Liu, and Igor Gilitschenski · 2023
Later among the works it cites.
Hidanet: Rgb-d salient object detection via hierarchical depth awareness
Zongwei Wu, Guillaume Allibert, Fabrice Meriaudeau, Chao Ma, and Cédric Demonceaux · 2023
Later among the works it cites.
Cmx: Cross-modal fusion for rgb-x semantic segmentation with transformers
Jiaming Zhang, Huayao Liu, Kailun Yang, Xinxin Hu, Ruiping Liu, and Rainer Stiefelhagen · 2023
Later among the works it cites.
Llama-adapter: Efficient fine-tuning of language models with zero-init attention
Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, Peng Gao, and Yu Qiao · 2023
Later among the works it cites.
Meta-transformer: A unified framework for multimodal learning
Yiyuan Zhang, Kaixiong Gong, Kaipeng Zhang, Hongsheng Li, Yu Qiao, Wanli Ouyang, and Xiangyu Yue · 2023
Later among the works it cites.
Arkittrack: a new diverse dataset for tracking using mobile rgb-d data
Haojie Zhao, Junsong Chen, Lijun Wang, and Huchuan Lu · 2023
Later among the works it cites.
Cvt-slr: Contrastive visual-textual transformation for sign language recognition with variational alignment
Jiangbin Zheng, Yile Wang, Cheng Tan, Siyuan Li, Ge Wang, Jun Xia, Yidong Chen, and Stan Z Li · 2023
Later among the works it cites.
E-clip: Towards label-efficient event-based open-world understanding by clip
Jiazhou Zhou, Xu Zheng, Yuanhuiyi Lyu, and Lin Wang · 2023
Later among the works it cites.
Bin Zhu, Bin Lin, Munan Ning, Yang Yan, Jiaxi Cui, HongFa Wang, Yatian Pang, Wenhao Jiang, Junwu Zhang, Zongwei Li, et al · 2023
Later among the works it cites.
Fourier prompt tuning for modality-incomplete scene segmentation
Ruiping Liu, Jiaming Zhang, Kunyu Peng, Yufan Chen, Ke Cao, Junwei Zheng, M Saquib Sarfraz, Kailun Yang, and Rainer Stiefelhagen · 2024
Closest in time.
Image anything: Towards reasoning-coherent and training-free multi-modal image generation
Yuanhuiyi Lyu, Xu Zheng, and Lin Wang · 2024
Closest in time.
Unibind: Llm-augmented unified and balanced representation space to bind them all
Yuanhuiyi Lyu, Xu Zheng, Jiazhou Zhou, and Lin Wang · 2024
Closest in time.
Any-to-any generation via composable diffusion
Zineng Tang, Ziyi Yang, Chenguang Zhu, Michael Zeng, and Mohit Bansal · 2024
Closest in time.
Binding touch to everything: Learning unified multimodal tactile representations
Fengyu Yang, Chao Feng, Ziyang Chen, Hyoungseob Park, Daniel Wang, Yiming Dou, Ziyao Zeng, Xien Chen, Rit Gangopadhyay, Andrew Owens, et al · 2024
Closest in time.
Learning unseen modality interaction
Yunhua Zhang, Hazel Doughty, and Cees Snoek · 2024
Closest in time.
Eventdance: Unsupervised source-free cross-modal adaptation for event-based object recognition
Xu Zheng and Lin Wang · 2024
Closest in time.