Fetching the paper…
Reading the bibliography…
This survey provides a comprehensive overview of recent advances in multimodal alignment and fusion within the field of machine learning, driven by the increasing availability and diversity of data modalities such as text, images, audio, and video.
Relations between two sets of variates
Harold Hotelling · 1936
Earlier work this paper cites.
An overview of sequence comparison: Time warps, string edits, and macromolecules
Joseph B. Kruskal · 1983
Earlier work this paper cites.
A kernel method for canonical correlation analysis
S. Akaho · 2001
Earlier work this paper cites.
Nonlinear feature extraction using generalized canonical correlation analysis
T. Melzer, M. Reiter, and H. Bischof · 2001
Earlier work this paper cites.
Kernel independent component analysis
F. R. Bach and M. I. Jordan · 2002
Earlier work this paper cites.
Canonical correlation analysis: An overview with application to learning methods
D. R. Hardoon, S. Szedmak, and J. Shawe-Taylor · 2004
Earlier work this paper cites.
Early versus late fusion in semantic video analysis
Cees G. M. Snoek, Marcel Worring, and Arnold W. M. Smeulders · 2005
Earlier work this paper cites.
Distributed fusion in sensor networks: a graphical models perspective
Müjdat Çetin, Lei Chen, John W. Fisher III, Alexander T. Ihler, Randolph L. Moses, Martin J. Wainwright, and Alan S. Willsky · 2006
Earlier work this paper cites.
Classifier fusion for svm-based multimedia semantic indexing
S. Ayache, Georges Quénot, and Jérôme Gensel · 2007
Earlier work this paper cites.
Dynamic time warping
Unknown · 2007
Earlier work this paper cites.
Functional learning of kernels for information fusion purposes
Alberto Muñoz and Javier González · 2008
Earlier work this paper cites.
Feature fusion hierarchies for gender classification
F. Scalzo, George Bebis, Mircea Nicolescu, Leandro A. Loss, and A. Tavakkoli · 2008
Earlier work this paper cites.
Model level fusion of edge histogram descriptors and gabor wavelets for landmine detection with ground penetrating radar
Oualid Missaoui, Hichem Frigui, and Paul D. Gader · 2010
Earlier work this paper cites.
A hierarchical feature fusion framework for adaptive visual tracking
Alexandros Makris, Dimitrios I. Kosmopoulos, Stavros J. Perantonis, and Sergios Theodoridis · 2011
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Vicente Ordonez, Girish Kulkarni, and Tamara Berg · 2011
Earlier work this paper cites.
Kernel-based data fusion improves the drug-protein interaction prediction
Yong-Cui Wang, Chunhua Zhang, Naiyang Deng, and Yong Wang · 2011
Earlier work this paper cites.
Fusing information from multifidelity computer models of physical systems
Douglas L. Allaire and Karen E. Willcox · 2012
Earlier work this paper cites.
Multi-aspect candidates for repositioning: data fusion methods using heterogeneous information sources
Adam Arany, Bence Bolgár, Balázs Balogh, Péter Antal, and Péter Mátyus · 2012
Earlier work this paper cites.
Deap: A database for emotion analysis ;using physiological signals
Sander Koelstra, Christian Muhl, Mohammad Soleymani, Jong-Seok Lee, Ashkan Yazdani, Touradj Ebrahimi, Thierry Pun, Anton Nijholt, and Ioannis Patras · 2012
Earlier work this paper cites.
Graphalignment: Bayesian pairwise alignment of biological networks
M. Kolář, J. Meier, V. Mustonen, et al · 2012
Earlier work this paper cites.
Introducing a new benchmarked dataset for activity monitoring
Attila Reiss and Didier Stricker · 2012
Earlier work this paper cites.
Multimodal learning with deep boltzmann machines
Nitish Srivastava and Ruslan Salakhutdinov · 2012
Earlier work this paper cites.
Kernel cross-modal factor analysis for information fusion with application to bimodal emotion recognition
Yongjin Wang, Ling Guan, and Anastasios N. Venetsanopoulos · 2012
Earlier work this paper cites.
Deep canonical correlation analysis
Galen Andrew, Raman Arora, Jeff Bilmes, and Karen Livescu · 2013
Earlier work this paper cites.
mhealthdroid: A novel framework for agile development of mobile health applications
Oresti Banos, Rafael Garcia, Juan A. Holgado-Terriza, Miguel Damas, Hector Pomares, Ignacio Rojas, Alejandro Saez, and Claudia Villalonga · 2014
Earlier work this paper cites.
Visualization of graphical information fusion results
Erik Blasch, Georgiy M. Levchuk, Gennady Staskevich, Dustin Burke, and Alex Aved · 2014
Earlier work this paper cites.
Majority vote of diverse classifiers for late fusion
Emilie Morvant, Amaury Habrard, and Stéphane Ayache · 2014
Earlier work this paper cites.
Im2text and text2im: Associating images and texts for cross-modal retrieval
Y. Verma and C. V. Jawahar · 2014
Earlier work this paper cites.
Multimodal data fusion in text-image heterogeneous graph for social media recommendation
Yu Xiong, Daling Wang, Yifei Zhang, Shi Feng, and Guoren Wang · 2014
Earlier work this paper cites.
Segnet: A deep convolutional encoder-decoder architecture for image segmentation
Vijay Badrinarayanan, Alex Kendall, and Roberto Cipolla · 2015
Earlier work this paper cites.
Learning to combine local models for facial action unit detection
Shashank Jaiswal, Brais Martínez, and Michel F. Valstar · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Fei-Fei Li · 2015
Earlier work this paper cites.
Microsoft coco: Common objects in context, 2015
Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár · 2015
Earlier work this paper cites.
Nonlinear graph fusion for multi-modal classification of alzheimer’s disease
Tong Tong, Katherine R. Gray, Qinquan Gao, Liang Chen, and Daniel Rueckert · 2015
Earlier work this paper cites.
Kernel-based sensor fusion with application to audio-visual voice activity detection
David Dov, Ronen Talmon, and Israel Cohen · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations, 2016
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Fei-Fei Li · 2016
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models, 2016
Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2016
Earlier work this paper cites.
Yfcc100m: the new data in multimedia research
Bart Thomee, David A. Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li · 2016
Earlier work this paper cites.
Multimodal image alignment via linear mapping between feature modalities
Yanyun Jiang, Yuanjie Zheng, Sujuan Hou, Yuchou Chang, and J. Gee · 2017
Earlier work this paper cites.
Applications of deep learning and reinforcement learning to biological data
M. Mahmud, M.S. Kaiser, A. Hussain, and Stefano Vassanelli · 2017
Earlier work this paper cites.
Multimodal network alignment
Huda Nassar and David Gleich · 2017
Earlier work this paper cites.
Multi-modal classification of alzheimer’s disease using nonlinear graph fusion
Tong Tong, Katherine R. Gray, Qinquan Gao, Liang Chen, and Daniel Rueckert · 2017
Earlier work this paper cites.
Tensor fusion network for multimodal sentiment analysis
Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency · 2017
Earlier work this paper cites.
Set cross entropy: Likelihood-based permutation invariant loss function for probability distributions, 2018
Masataro Asai · 2018
Earlier work this paper cites.
Multimodal machine learning: A survey and taxonomy
Tadas Baltrusaitis, Chaitanya Ahuja, and Louis-Philippe Morency · 2018
Earlier work this paper cites.
Densefuse: A fusion approach to infrared and visible images
Hui Li and Xiaojun Wu · 2018
Earlier work this paper cites.
Efficient low-rank multimodal fusion with modality-specific factors
Zhun Liu, Ying Shen, Varun Bharadhwaj Lakshminarasimhan, Paul Pu Liang, AmirAli Bagher Zadeh, and Louis-Philippe Morency · 2018
Earlier work this paper cites.
Design of a low-level radar and time-of-flight sensor fusion framework
Josef Steinbaeck, Christian Steger, Gerald Holweg, and Norbert Druml · 2018
Earlier work this paper cites.
Research of spatial alignment techniques for multimodal image fusion
A. Akhmerov, A. Vasilev, and A.V. Vasileva · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks, 2019
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Earlier work this paper cites.
Divide, conquer and combine: Hierarchical feature fusion network with local and global perspectives for multimodal affective computing
Sijie Mai, Haifeng Hu, and Songlong Xing · 2019
Earlier work this paper cites.
Audio-visual emotion recognition in video clips
Fatemeh Noroozi, Marina Marjanovic, Angelina Njegus, Sergio Escalera, and Gholamreza Anbarjafari · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Towards raw sensor fusion in 3d object detection
András Rövid and Viktor Remeli · 2019
Earlier work this paper cites.
Fast and robust dynamic hand gesture recognition via key frames extraction and feature fusion
Hao Tang, Hong Liu, Wei Xiao, and Nicu Sebe · 2019
Earlier work this paper cites.
Multi-channel attention selection gan with cascaded semantic guidance for cross-view image translation
Hao Tang, Dan Xu, Nicu Sebe, Yanzhi Wang, Jason J Corso, and Yan Yan · 2019
Earlier work this paper cites.
Multimodal deep representation learning for video classification
Haiman Tian, Yudong Tao, Samira Pouyanfar, Shu-Ching Chen, and Mei-Ling Shyu · 2019
Earlier work this paper cites.
Learning factorized multimodal representations, 2019
Yao-Hung Hubert Tsai, Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
Multimodal generative models for compositional representation learning, 2019
Mike Wu and Noah Goodman · 2019
Earlier work this paper cites.
Hgmf: Heterogeneous graph-based fusion for multimodal data with incompleteness
Jiayi Chen and Aidong Zhang · 2020
Earlier work this paper cites.
Uniter: Universal image-text representation learning, 2020
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2020
Earlier work this paper cites.
Multimodal mri synthesis using unified generative adversarial networks
X. Dai, Y. Lei, Y. Fu, W. Curran, T. Liu, H. Mao, and Xiaofeng Yang · 2020
Earlier work this paper cites.
Sensor fusion of camera and lidar raw data for vehicle detection
Gokulesh Danapal, Giovanni A. Santos, João Paulo C. L. da Costa, Bruno J. G. Praciano, and Gabriel P. M. Pinheiro · 2020
Earlier work this paper cites.
Multi-modal transformer for video retrieval
Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid · 2020
Earlier work this paper cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks, 2020
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, Yejin Choi, and Jianfeng Gao · 2020
Earlier work this paper cites.
Hierarchical feature fusion network for salient object detection
Xuelong Li, Dawei Song, and Yongsheng Dong · 2020
Earlier work this paper cites.
Modality to modality translation: An adversarial representation learning and graph fusion network for multimodal fusion, 2020
Sijie Mai, Haifeng Hu, and Songlong Xing · 2020
Earlier work this paper cites.
Weakly supervised representation learning for audio-visual scene analysis
Sanjeel Parekh, Slim Essid, Alexey Ozerov, Ngoc Q. K. Duong, Patrick Pérez, and Gaël Richard · 2020
Earlier work this paper cites.
Local class-specific and global image-level generative adversarial networks for semantic-guided scene generation
Hao Tang, Dan Xu, Yan Yan, Philip HS Torr, and Nicu Sebe · 2020
Earlier work this paper cites.
Guided deep decoder: Unsupervised image pair fusion
Tatsumi Uezato, Danfeng Hong, Naoto Yokoya, and Wei He · 2020
Earlier work this paper cites.
Scene graph-based semantic alignment for multimodal tasks
Wei Xiong, Yifan Zhang, and Wei Li · 2020
Earlier work this paper cites.
Mitigating biases in multimodal personality assessment
Shen Yan, Di Huang, and Mohammad Soleymani · 2020
Earlier work this paper cites.
CH-SIMS: A Chinese multimodal sentiment analysis dataset with fine-grained annotation of modality
Wenmeng Yu, Hua Xu, Fanyang Meng, Yilin Zhu, Yixiao Ma, Jiele Wu, Jiyun Zou, and Kaicheng Yang · 2020
Earlier work this paper cites.
Multimodal intelligence: Representation learning, information fusion, and applications
Chao Zhang, Zichao Yang, Xiaodong He, and Li Deng · 2020
Earlier work this paper cites.
Harnessing multimodal data integration to advance precision oncology
K. Boehm, P. Khosravi, R. Vanguri, Jianjiong Gao, and S. Shah · 2021
Earlier work this paper cites.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts, 2021
Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut · 2021
Earlier work this paper cites.
Redcaps: web-curated image-text data created by the people, for the people, 2021
Karan Desai, Gaurav Kaul, Zubin Aysola, and Justin Johnson · 2021
Earlier work this paper cites.
Sparse fusion for multimodal transformers, 2021
Yi Ding, Alex Rich, Mason Wang, Noah Stier, Matthew Turk, Pradeep Sen, and Tobias Höllerer · 2021
Earlier work this paper cites.
Audio-visual event localization via recursive fusion by joint co-attention
Bin Duan, Hao Tang, Wei Wang, Ziliang Zong, Guowei Yang, and Yan Yan · 2021
Earlier work this paper cites.
Globalizing BERT-based transformer architectures for long document summarization
Quentin Grail, Julien Perez, and Eric Gaussier · 2021
Earlier work this paper cites.
Improving multimodal fusion with hierarchical mutual information maximization for multimodal sentiment analysis
Wei Han, Hui Chen, and Soujanya Poria · 2021
Earlier work this paper cites.
Perceiver: General perception with iterative attention, 2021
Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and Joao Carreira · 2021
Earlier work this paper cites.
Scaling up visual and vision-language representation learning with noisy text supervision, 2021
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yunhsuan Sung, Zhen Li, and Tom Duerig · 2021
Earlier work this paper cites.
Align before fuse: Vision and language representation learning with momentum distillation
Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi · 2021
Cited alongside, same era.
Vx2text: End-to-end learning of video-based text generation from multimodal inputs
Xudong Lin, Gedas Bertasius, Jue Wang, Shih-Fu Chang, Devi Parikh, and Lorenzo Torresani · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning
Krishna Srinivasan, Karthik Raman, Jiecao Chen, Michael Bendersky, and Marc Najork · 2021
Cited alongside, same era.
The multimodal sentiment analysis in car reviews (muse-car) dataset: Collection, insights and improvements, 2021
Lukas Stappen, Alice Baird, Lea Schumann, and Björn Schuller · 2021
Bafn: Bi-direction attention based fusion network for multimodal sentiment analysis
Jiajia Tang, Dongjun Liu, Xuanyu Jin, Yong Peng, Qianchuan Zhao, Yu Ding, and Wanzeng Kong · 2023
Later among the works it cites.
Uit-saviors at medvqa-gi 2023: Improving multimodal learning with image enhancement for gastrointestinal visual question answering, 2023
T. M. Thai, A. T. Vo, Hao K. Tieu, Linh Bui, and T. Nguyen · 2023
Later among the works it cites.
Villa: Fine-grained vision-language representation learning from real-world data, 2023
Maya Varma, Jean-Benoit Delbrouck, Sarah Hooper, Akshay Chaudhari, and Curtis Langlotz · 2023
Later among the works it cites.
Few-shot in-context imitation learning via implicit graph alignment
Vitalis Vosylius and Edward Johns · 2023
Later among the works it cites.
Too large; data reduction for vision-language pre-training, 2023
Alex Jinpeng Wang, Kevin Qinghong Lin, David Junhao Zhang, Stan Weixian Lei, and Mike Zheng Shou · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Attentiongan: Unpaired image-to-image translation using attention-guided generative adversarial networks
Hao Tang, Hong Liu, Dan Xu, Philip HS Torr, and Nicu Sebe · 2021
Cited alongside, same era.
Graph-based multimodal sequential embedding for sign language translation
Shengeng Tang, Dixin Guo, Rui Hong, and Min Wang · 2021
Cited alongside, same era.
Decision-level data fusion in quality control and predictive maintenance
Yupeng Wei, Dazhong Wu, and Janis P. Terpenny · 2021
Cited alongside, same era.
Towards user friendly medication mapping using entity-boosted two-tower neural network
S. Yuan, P. Bhatia, B. Celikkaya, H. Liu, and K. Choi · 2021
Cited alongside, same era.
Vinvl: Revisiting visual representations in vision-language models, 2021
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao · 2021
Cited alongside, same era.
Multimodal deep fusion for image question answering
Weifeng Zhang, Jing Yu, Yuxia Wang, and Wei Wang · 2021
Cited alongside, same era.
A token-wise graph-based framework for multimodal named entity recognition
Zhiwei Zhang, Wenyu Mai, Heng Xiong, and Cheng Wu · 2021
Cited alongside, same era.
Mutually beneficial transformer for multimodal data fusion
Jinping Wang and Xiaojun Tan · 2023
Later among the works it cites.
One-peace: Exploring one general representation model toward unlimited modalities, 2023
Peng Wang, Shijie Wang, Junyang Lin, Shuai Bai, Xiaohuan Zhou, Jingren Zhou, Xinggang Wang, and Chang Zhou · 2023
Later among the works it cites.
Bridgetower: Building bridges between encoders in vision-language representation learning
X. Xu, C. Wu, S. Rosenman, V. Lal, and W. Che · 2023
Later among the works it cites.
Dynamic multimodal fusion
Zihui Xue and Radu Marculescu · 2023
Later among the works it cites.
Videochat: Conversational agents in video understanding
H. Yang and S. Li · 2023
Later among the works it cites.
Macsa: A multimodal aspect-category sentiment analysis dataset with multimodal fine-grained aligned annotations
Haoyan Yang, Yifan Wu, Zhenyu Si, Yijun Zhao, Jinfeng Liu, and Bing Qin · 2023
Later among the works it cites.
When and why vision-language models behave like bags-of-words, and what to do about it?, 2023
Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky, and James Zou · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training, 2023
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer · 2023
Later among the works it cites.
Neural attention: Enhancing qkv calculation in self-attention mechanism with neural networks, 2023
Muhan Zhang · 2023
Later among the works it cites.
Internlm-xcomposer: A vision-language large model for advanced text-image comprehension and composition, 2023
Pan Zhang, Xiaoyi Dong, Bin Wang, Yuhang Cao, Chao Xu, Linke Ouyang, Zhiyuan Zhao, Haodong Duan, Songyang Zhang, Shuangrui Ding, Wenwei Zhang, Hang Yan, Xinyue Zhang, Wei Li, Jingwen Li, Kai Chen, Conghui He, Xingcheng Zhang, Yu Qiao, Dahua Lin, and Jiaqi Wang · 2023
Later among the works it cites.
Bimodal fusion network with multi-head attention for multimodal sentiment analysis
Rui Zhang, Chengrong Xue, Qingfu Qi, Liyuan Lin, Jing Zhang, and Lun Zhang · 2023
Later among the works it cites.
Enlighten-your-voice: When multimodal meets zero-shot low-light image enhancement
Xiaofeng Zhang, Zishan Xu, Hao Tang, Chaochen Gu, Wei Chen, Shanying Zhu, and Xinping Guan · 2023
Later among the works it cites.
Vision + language applications: A survey
Yutong Zhou and Nobutaka Shimada · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny · 2023
Later among the works it cites.
Svp-t: A shape-level variable-position transformer for multivariate time series classification
Rui Zuo, Guoqing Li, Bongshin Choi, Sourav Bhowmick, Daphne N. yin Mah, and Gary L. Wong · 2023
Later among the works it cites.
4m-21: An any-to-any vision model for tens of tasks and modalities, 2024
Roman Bachmann, Oğuzhan Fatih Kar, David Mizrahi, Ali Garjani, Mingfei Gao, David Griffiths, Jiaming Hu, Afshin Dehghan, and Amir Zamir · 2024
Closest in time.
Condition-aware multimodal fusion for robust semantic perception of driving scenes, 2024
Tim Broedermann, Christos Sakaridis, Yuqian Fu, and Luc Van Gool · 2024
Closest in time.
Mix-tower: Light visual question answering framework based on exclusive self-attention mechanism
D. Chen, J. Chen, L. Yang, and F. Shang · 2024
Closest in time.
Comkd-clip: Comprehensive knowledge distillation for contrastive language-image pre-traning model, 2024
Yifan Chen, Xiaozhen Qiao, Zhe Sun, and Xuelong Li · 2024
Closest in time.
Multimodal representation learning for tourism recommendation with two-tower architecture
Y. Cui, S. Liang, and YY Zhang · 2024
Closest in time.
Moshi: a speech-text foundation model for real-time dialogue, 2024
Alexandre Défossez, Laurent Mazaré, Manu Orsini, Amélie Royer, Patrick Pérez, Hervé Jégou, Edouard Grave, and Neil Zeghidour · 2024
Closest in time.
Onellm: One framework to align all modalities with language
Han et al · 2024
Closest in time.
Domain aligned clip for few-shot classification
Muhammad Waleed Gondal, Jochen Gast, Inigo Alonso Ruiz, Richard Droste, Tommaso Macri, Suren Kumar, and Luitpold Staudigl · 2024
Closest in time.
Geminifusion: Efficient pixel-wise multimodal fusion for vision transformer, 2024
Ding Jia, Jianyuan Guo, Kai Han, Han Wu, Chao Zhang, Chang Xu, and Xinghao Chen · 2024
Closest in time.
Diffusion-driven gan inversion for multi-modal face image generation
Jihyun Kim, Changjae Oh, Hoseok Do, Soohyun Kim, and Kwanghoon Sohn · 2024
Closest in time.
Hype: Hyperbolic entailment filtering for underspecified images and texts
Wonjae Kim, Sanghyuk Chun, Taekyung Kim, Dongyoon Han, and Sangdoo Yun · 2024
Closest in time.
Gs-clip: Gaussian splatting for contrastive language-image-3d pretraining from real-world data, 2024
Haoyuan Li, Yanpeng Zhou, Yihan Zeng, Hang Xu, and Xiaodan Liang · 2024
Closest in time.
Rethinking transformer for long contextual histopathology whole slide image analysis
Honglin Li, Yunlong Zhang, Pingyi Chen, Zhongyi Shui, Chenglu Zhu, and Lin Yang · 2024
Closest in time.
TextBind: Multi-turn interleaved multimodal instruction-following in the wild
Huayang Li, Siheng Li, Deng Cai, Longyue Wang, Lemao Liu, Taro Watanabe, Yujiu Yang, and Shuming Shi · 2024
Closest in time.
Coupled mamba: Enhanced multi-modal fusion with coupled state space model, 2024
Wenbing Li, Hang Zhou, Junqing Yu, Zikai Song, and Wei Yang · 2024
Closest in time.
Data processing techniques for modern multimodal models
Yinheng Li, Han Ding, and Hang Chen · 2024
Closest in time.
Foundations & trends in multimodal machine learning: Principles, challenges, and open questions
Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency · 2024
Closest in time.
Vila: On pre-training for visual language models, 2024
Ji Lin, Hongxu Yin, Wei Ping, Yao Lu, Pavlo Molchanov, Andrew Tao, Huizi Mao, Jan Kautz, Mohammad Shoeybi, and Song Han · 2024
Closest in time.
St-align: A multimodal foundation model for image-gene alignment in spatial transcriptomics, 2024
Yuxiang Lin, Ling Luo, Ying Chen, Xushi Zhang, Zihui Wang, Wenxian Yang, Mengsha Tong, and Rongshan Yu · 2024
Closest in time.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2024
Closest in time.
Learning modality knowledge alignment for cross-modality transfer
Wenxuan Ma, Shuang Li, Lincan Cai, and Jingxuan Kang · 2024
Closest in time.
Sycoca: Symmetrizing contrastive captioners with attentive masking for multimodal alignment, 2024
Ziping Ma, Furong Xu, Jian Liu, Ming Yang, and Qingpei Guo · 2024
Closest in time.
A survey on multimodal wearable sensor-based human action recognition
Jianyuan Ni, Hao Tang, Syed Tousiful Haque, Yan Yan, and Anne HH Ngu · 2024
Closest in time.
Abftnet: An efficient transformer network with alignment before fusion for multimodal automatic modulation recognition
Meng Ning, Fan Zhou, Wei Wang, Shaoqiang Wang, Peiying Zhang, and Jian Wang · 2024
Closest in time.
Dinov2: Learning robust visual features without supervision, 2024
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, and et al · 2024
Closest in time.
Survey of large multimodal model datasets, application categories and taxonomy, 2024
Priyaranjan Pattnayak, Hitesh Laxmichand Patel, Bhargava Kumar, Amit Agarwal, Ishan Banerjee, Srikant Panda, and Tejaswini Kumar · 2024
Closest in time.
Towards ethical multimodal systems, 2024
Alexis Roger, Esma Aïmeur, and Irina Rish · 2024
Closest in time.
Set-clip: Exploring aligned semantic from low-alignment multimodal data through a distribution view, 2024
Zijia Song, Zelin Zang, Yelin Wang, Guozheng Yang, Kaicheng yu, Wanyu Chen, Miaoyu Wang, and Stan Z. Li · 2024
Closest in time.
Graph transformer gans with graph masked modeling for architectural layout generation
Hao Tang, Ling Shao, Nicu Sebe, and Luc Van Gool · 2024
Closest in time.
COLD fusion: Calibrated and ordinal latent distribution fusion for uncertainty-aware multimodal emotion recognition
Mani Kumar Tellamekala, Shahin Amiriparian, Björn W. Schuller, Elisabeth André, Timo Giesbrecht, and Michel Valstar · 2024
Closest in time.
Eyes wide shut? exploring the visual shortcomings of multimodal llms
Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, and Saining Xie · 2024
Closest in time.
I can listen but cannot read: An evaluation of two-tower multimodal systems for instrument recognition, 2024
Yannis Vasilakis, Rachel Bittner, and Johan Pauwels · 2024
Closest in time.
Data-efficient multimodal fusion on a single gpu, 2024
Noël Vouitsis, Zhaoyan Liu, Satya Krishna Gorti, Valentin Villecroze, Jesse C. Cresswell, Guangwei Yu, Gabriel Loaiza-Ganem, and Maksims Volkovs · 2024
Closest in time.
Cross-modal feature alignment and fusion for composed image retrieval
Yongquan Wan, Wenhai Wang, Guobing Zou, and Bofeng Zhang · 2024
Closest in time.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin · 2024
Closest in time.
Vila: Efficient video-language alignment for video question answering, 2024
Xijun Wang, Junbang Liang, Chun-Kai Wang, Kenan Deng, Yu Lou, Ming Lin, and Shan Yang · 2024
Closest in time.
Multimodal reranking for knowledge-intensive visual question answering, 2024
Haoyang Wen, Honglei Zhuang, Hamed Zamani, Alexander Hauptmann, and Michael Bendersky · 2024
Closest in time.
Mfeclip: Clip with mapping-fusion embedding for text-guided image editing
Fei Wu, Yongheng Ma, Hao Jin, Xiao-Yuan Jing, and Guo-Ping Jiang · 2024
Closest in time.
Dmf-gan: Deep multimodal fusion generative adversarial networks for text-to-image synthesis
Bing Yang, Xueqin Xiang, Wangzeng Kong, Jianhai Zhang, and Yong Peng · 2024
Closest in time.
Yi: Open foundation models by 01.ai, 2024
Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, Kaidong Yu, Peng Liu, Qiang Liu, Shawn Yue, Senbin Yang, Shiming Yang, Tao Yu, Wen Xie, Wenhao Huang, Xiaohui Hu, Xiaoyi Ren, Xinyao Niu, Pengcheng Nie, Yuchi Xu, Yudong Liu, Yue Wang, Yuxuan Cai, Zhenyu Gu, Zhiyuan Liu, and Zonghong Dai · 2024
Closest in time.
Capsfusion: Rethinking image-text data at scale, 2024
Qiying Yu, Quan Sun, Xiaosong Zhang, Yufeng Cui, Fan Zhang, Yue Cao, Xinlong Wang, and Jingjing Liu · 2024
Closest in time.
Contextual object detection with multimodal large language models, 2024
Yuhang Zang, Wei Li, Jun Han, Kaiyang Zhou, and Chen Change Loy · 2024
Closest in time.
MM-LLMs: Recent advances in MultiModal large language models
Duzhen Zhang, Yahan Yu, Jiahua Dong, Chenxing Li, Dan Su, Chenhui Chu, and Dong Yu · 2024
Closest in time.
Vision-language models for vision tasks: A survey
Jingyi Zhang, Jiaxing Huang, Sheng Jin, and Shijian Lu · 2024
Closest in time.
Internlm-xcomposer-2.5: A versatile large vision language model supporting long-contextual input and output, 2024
Pan Zhang, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Rui Qian, Lin Chen, Qipeng Guo, Haodong Duan, Bin Wang, Linke Ouyang, Songyang Zhang, Wenwei Zhang, Yining Li, Yang Gao, Peng Sun, Xinyue Zhang, Wei Li, Jingwen Li, Wenhai Wang, Hang Yan, Conghui He, Xingcheng Zhang, Kai Chen, Jifeng Dai, Yu Qiao, Dahua Lin, and Jiaqi Wang · 2024
Closest in time.
Genderalign: An alignment dataset for mitigating gender bias in large language models, 2024
Tao Zhang, Ziqian Zeng, Yuxiang Xiao, Huiping Zhuang, Cen Chen, James Foulds, and Shimei Pan · 2024
Closest in time.
From redundancy to relevance: Enhancing explainability in multimodal large language models
Xiaofeng Zhang, Chen Shen, Xiaosong Yuan, Shaotian Yan, Liang Xie, Wenxiao Wang, Chaochen Gu, Hao Tang, and Jieping Ye · 2024
Closest in time.
Understanding unimodal bias in multimodal deep linear networks, 2024
Yedi Zhang, Peter E. Latham, and Andrew Saxe · 2024
Closest in time.
Risurconv: Rotation invariant surface attention-augmented convolutions for 3d point cloud classification and segmentation, 2024
Zhiyuan Zhang, Licheng Yang, and Zhiyu Xiang · 2024
Closest in time.
Rs5m and georsclip: A large-scale vision- language dataset and a large vision-language model for remote sensing
Zilun Zhang, Tiancheng Zhao, Yulong Guo, and Jianwei Yin · 2024
Closest in time.
Deep multimodal data fusion
Fei Zhao, Chengcui Zhang, and Baocheng Geng · 2024
Closest in time.
Deep multimodal learning with vision, audio, and text: Challenges and innovations
Lihong Zhao and Huan Wang · 2024
Closest in time.
Generative ai for vision: A comprehensive study of frameworks and applications, 2025
Fouad Bousetouane · 2025
Closest in time.
Large language models and large multimodal models in medical imaging: A primer for physicians
Tyler J. Bradshaw, Xin Tie, Joshua Warner, Junjie Hu, Quanzheng Li, and Xiang Li · 2025
Closest in time.
Generative ai models: Theoretical foundations and algorithmic practices
Yongnian Cao, Xuechun Yang, and Rui Sun · 2025
Closest in time.
Autovit: Achieving real-time vision transformers on mobile via latency-aware coarse-to-fine search
Zhenglun Kong, Dongkuan Xu, Zhengang Li, Peiyan Dong, Hao Tang, Yanzhi Wang, and Subhabrata Mukherjee · 2025
Closest in time.
Mulfs-cap: Multimodal fusion-supervised cross-modality alignment perception for unregistered infrared-visible image fusion
Huafeng Li, Zengyi Yang, Yafei Zhang, Wei Jia, Zhengtao Yu, and Yu Liu · 2025
Closest in time.
A survey of state of the art large vision language models: Alignment, benchmark, evaluations and challenges, 2025
Zongxia Li, Xiyang Wu, Hongyang Du, Fuxiao Liu, Huy Nghiem, and Guangyao Shi · 2025
Closest in time.
Enhanced multi-scale cross-attention for person image generation
Hao Tang, Ling Shao, Nicu Sebe, and Luc Van Gool · 2025
Closest in time.
TTTFusion: A Test-Time Training-Based Strategy for Multimodal Medical Image Fusion in Surgical Robots
Qinhua Xie and Hao Tang · 2025
Closest in time.
Qwen2.5-omni technical report, 2025
Jin Xu, Zhifang Guo, Jinzheng He, Hangrui Hu, Ting He, Shuai Bai, Keqin Chen, Jialin Wang, Yang Fan, Kai Dang, Bin Zhang, Xiong Wang, Yunfei Chu, and Junyang Lin · 2025
Closest in time.
Q-tempfusion: Quantization-aware temporal multi-sensor fusion on bird’s-eye view representation
Pinrui Yu, Zhenglun Kong, Pu Zhao, Peiyan Dong, Hao Tang, Fei Sun, Xue Lin, and Yanzhi Wang · 2025
Closest in time.
Llava-video: Video instruction tuning with synthetic data, 2025
Yuanhan Zhang, Jinming Wu, Wei Li, Bo Li, Zejun Ma, Ziwei Liu, and Chunyuan Li · 2025
Closest in time.