Fetching the paper…
Reading the bibliography…
In the realm of Medical Visual Language Models (Med-VLMs), the quest for universal efficient fine-tuning mechanisms remains paramount, especially given researchers in interdisciplinary fields are often extremely short of training resources, yet largely unexplored.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Ryder, et al · 1901
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Lee, et al · 2018
Earlier work this paper cites.
A dataset of clinically generated visual questions and answers about radiology images
Jason J Lau, Gayen, et al · 2018
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP. In International Conference on Machine Learning . PMLR, 2790–2799
Neil Houlsby, Andrei Giurgiu, Jastrzebski, et al · 2019
Earlier work this paper cites.
MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports
Alistair EW Johnson, Tom J Pollard, Seth J Berkowitz, Nathaniel R Greenbaum, Matthew P Lungren, Chih-ying Deng, Roger G Mark, and Steven Horng. 2019 · 2019
Earlier work this paper cites.
Overcoming Data Limitation in Medical Visual Question Answering. In MICCAI . Cham, 522–530
Binh D. Nguyen, Thanh-Toan Do, Binh X Nguyen, et al · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
MedICaT: A Dataset of Medical Images, Captions, and Textual References. In Findings of EMNLP
Sachin Mehta Sanjay Subramanian, Lucy Lu Wang et al · 2020
Earlier work this paper cites.
Multiple Meta-model Quantifying for Medical Visual Question Answering. In MICCAI . Cham, 64–74
Tuong Do, Binh X. Nguyen, et al · 2021
Earlier work this paper cites.
Learning Associative Representation for Facial Expression Recognition. In IEEE International Conference on Image Processing (ICIP) . 889–893
Yangtao Du, Dingkang Yang, Peng Zhai, Mingchen Li, and Lihua Zhang. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Shen, et al · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Earlier work this paper cites.
Align before fuse: Vision and language representation learning with momentum distillation
Junnan Li, Ramprasaath Selvaraju, Gotmare, et al · 2021
Earlier work this paper cites.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang. 2021 · 2021
Earlier work this paper cites.
Contrastive Pre-training and Representation Distillation for Medical Visual Question Answering Based on Radiology Images. In MICCAI 2021 . Springer International Publishing, Cham, 210–220
Bo Liu, Li-Ming Zhan, and Xiao-Ming Wu. 2021a · 2021
Earlier work this paper cites.
Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering. In 2021 ISBI . 1650–1654
Bo Liu, Li-Ming Zhan, Xu, et al · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision. In International conference on machine learning . PMLR, 8748–8763
Alec Radford, Jong Wook Kim, Hallacy, et al · 2021
Cited alongside, same era.
Multi-modal masked autoencoders for medical vision-and-language pre-training. In MICCAI . Springer, 679–689
Zhihong Chen, Yuhao Du, Hu, et al · 2022
Cited alongside, same era.
VQAMix: Conditional Triplet Mixup for Medical Visual Question Answering
Haifan Gong, Guanqi Chen, Mao, et al · 2022
Cited alongside, same era.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In ICCV . 12888–12900
Junnan Li, Dongxu Li, Xiong, et al · 2022
Cited alongside, same era.
Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5227–5237
Yi-Lin Sung, Jaemin Cho, and Mohit Bansal. 2022 · 2022
The expressive power of low-rank adaptation
Yuchen Zeng and Kangwook Lee. 2023 · 2023
Later among the works it cites.
Llama-adapter: Efficient fine-tuning of language models with zero-init attention
Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, Peng Gao, and Yu Qiao. 2023 · 2023
Later among the works it cites.
Tuning LayerNorm in Attention: Towards efficient multi-modal llm finetuning
Bingchen Zhao, Haoqin Tu, Chen Wei, Jieru Mei, and Cihang Xie. 2023 · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning. arXiv 2023
W Dai, J Li, D Li, AMH Tiong, J Zhao, W Wang, B Li, P Fung, and S Hoi. [n. d.] · 2023
Cited alongside, same era.
Parameter-efficient model adaptation for vision transformers. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 37. 817–825
Xuehai He, Chunyuan Li, Pengchuan Zhang, Jianwei Yang, and Xin Eric Wang. 2023 · 2023
Cited alongside, same era.
Yuxuan Lei, Dingkang Yang, Mingcheng Li, Shunli Wang, Jiawei Chen, and Lihua Zhang. 2023 · 2023
Cited alongside, same era.
Llava-med: Training a large language-and-vision assistant for biomedicine in one day
Chunyuan Li, Cliff Wong, Zhang, et al · 2023
Cited alongside, same era.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023a · 2023
Cited alongside, same era.
Towards Robust Multimodal Sentiment Analysis under Uncertain Signal Missing
Mingcheng Li, Dingkang Yang, and Lihua Zhang. 2023c · 2023
Cited alongside, same era.
Haotian Liu, Chunyuan Li, Wu, et al · 2023
Cited alongside, same era.
Jiawei Chen, Yue Jiang, Dingkang Yang, Mingcheng Li, Jinjie Wei, Ziyun Qian, and Lihua Zhang. 2024a · 2024
Closest in time.
MISS: A Generative Pretraining and Finetuning Approach for Med-VQA
Jiawei Chen, Dingkang Yang, Yue Jiang, et al · 2024
Closest in time.
Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey
Zeyu Han, Chao Gao, Jinyang Liu, Sai Qian Zhang, et al · 2024
Closest in time.
A Unified Self-Distillation Framework for Multimodal Sentiment Analysis with Uncertain Missing Modalities. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI) , Vol. 38. 10074–10082
Mingcheng Li, Dingkang Yang, Yuxuan Lei, Shunli Wang, Shuaibing Wang, Liuzhen Su, Kun Yang, Yuzheng Wang, Mingyang Sun, and Lihua Zhang. 2024 · 2024
Closest in time.
DoRA: Weight-Decomposed Low-Rank Adaptation
Shih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov, Yu-Chiang Frank Wang, Kwang-Ting Cheng, and Min-Hung Chen. 2024 · 2024
Closest in time.
Parameter-efficient tuning of large-scale multimodal foundation model
Haixin Wang, Xinlong Yang, Jianlong Chang, Dian Jin, Jinan Sun, Shikun Zhang, Xiao Luo, and Qi Tian. 2024 · 2024
Closest in time.
Towards Multimodal Sentiment Analysis Debiasing via Bias Purification
Dingkang Yang, Mingcheng Li, Dongling Xiao, Yang Liu, Kun Yang, Zhaoyu Chen, Yuzheng Wang, Peng Zhai, Ke Li, and Lihua Zhang. 2024a · 2024
Closest in time.
Towards Multimodal Human Intention Understanding Debiasing via Subject-Deconfounding
Dingkang Yang, Dongling Xiao, Ke Li, Yuzheng Wang, Zhaoyu Chen, Jinjie Wei, and Lihua Zhang. 2024b · 2024
Closest in time.
Robust Emotion Recognition in Context Debiasing
Dingkang Yang, Kun Yang, Mingcheng Li, Shunli Wang, Shuaibing Wang, and Lihua Zhang. 2024c · 2024
Closest in time.
Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2097–2106
Xiaosong Wang, Yifan Peng, Le Lu, Zhiyong Lu, Mohammadhadi Bagheri, and Ronald M Summers. 2017 · 2097
Closest in time.
Contextual and Cross-Modal Interaction for Multi-Modal Speech Emotion Recognition
Dingkang Yang, Shuai Huang, Yang Liu, and Lihua Zhang. 2022b · 2097
Closest in time.