Fetching the paper…
Reading the bibliography…
This survey delves into the realm of Parameter-Efficient Fine-Tuning (PEFT) within the context of Foundation Models (FMs).
E. T. K. Sang and F. D. Meulder, “Introduction to the conll-2003 shared task: Language-independent named entity recognition,” in Conference on Computational Natural Language Learning , 2003
2003
Earlier work this paper cites.
K.-L. Liu, W.-J. Li, and M. Guo, “Emoticon smoothed language models for twitter sentiment analysis,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 26, no. 1, 2012, pp. 1678–1684
2012
Earlier work this paper cites.
X. Shi, Z. Chen, H. Wang, D. Y. Yeung, W.-K. Wong, and W. chun Woo, “Convolutional lstm network: A machine learning approach for precipitation nowcasting,” in NIPS , 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep residual networks,” in European Conference on Computer Vision , 2016
2016
Earlier work this paper cites.
D. Hendrycks and K. Gimpel, “Gaussian error linear units (gelus),” arXiv: Learning , 2016
2016
Earlier work this paper cites.
Y. Li, Y. Liang, and A. Risteski, “Recovery guarantee of weighted low-rank approximation via alternating minimization,” in International Conference on Machine Learning , 2016
2016
Earlier work this paper cites.
A. Vaswani, “Attention is all you need,” arXiv preprint arXiv:1706.03762 , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
S.-A. Rebuffi, H. Bilen, and A. Vedaldi, “Learning multiple visual domains with residual adapters,” in NIPS , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
C. Gardent, A. Shimorina, S. Narayan, and L. Perez-Beltrachini, “The webnlg challenge: Generating text from rdf data,” in International Conference on Natural Language Generation , 2017
2017
Earlier work this paper cites.
Y. Li, T. Ma, and H. Zhang, “Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations,” in Annual Conference Computational Learning Theory , 2017
2017
Earlier work this paper cites.
A. Radford, “Improving language understanding by generative pre-training,” OpenAI , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
X. Liu, Q. Chen, C. Deng, H.-J. Zeng, J. Chen, D. Li, and B. Tang, “Lcqmc:a large-scale chinese question matching corpus,” in International Conference on Computational Linguistics , 2018
2018
Earlier work this paper cites.
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman, “Glue: A multi-task benchmark and analysis platform for natural language understanding,” in BlackboxNLP@EMNLP , 2018
2018
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Conneau and G. Lample, “Cross-lingual language model pretraining,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. de Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, “Parameter-efficient transfer learning for nlp,” in International Conference on Machine Learning , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Gaier and D. R. Ha, “Weight agnostic neural networks,” in Neural Information Processing Systems , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
P.-E. Sarlin, D. DeTone, T. Malisiewicz, and A. Rabinovich, “Superglue: Learning feature matching with graph neural networks,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 4937–4946, 2019
2019
Earlier work this paper cites.
T. B. Brown, “Language models are few-shot learners,” arXiv preprint arXiv:2005.14165 , 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
M. Zhao, T. Lin, M. Jaggi, and H. Schütze, “Masking as an efficient alternative to finetuning for pretrained language models,” in Conference on Empirical Methods in Natural Language Processing , 2020
2020
Earlier work this paper cites.
D. Guo, A. M. Rush, and Y. Kim, “Parameter-efficient transfer learning with diff pruning,” in Annual Meeting of the Association for Computational Linguistics , 2020
2020
Earlier work this paper cites.
J. Pfeiffer, I. Vulic, I. Gurevych, and S. Ruder, “Mad-x: An adapter-based framework for multi-task cross-lingual transfer,” in Conference on Empirical Methods in Natural Language Processing , 2020
2020
Earlier work this paper cites.
A. Rücklé, G. Geigle, M. Glockner, T. Beck, J. Pfeiffer, N. Reimers, and I. Gurevych, “Adapterdrop: On the efficiency of adapters in transformers,” in Conference on Empirical Methods in Natural Language Processing , 2020
2020
Earlier work this paper cites.
T. Schick and H. Schütze, “Exploiting cloze-questions for few-shot text classification and natural language inference,” in Conference of the European Chapter of the Association for Computational Linguistics , 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
A. Aghajanyan, L. Zettlemoyer, and S. Gupta, “Intrinsic dimensionality explains the effectiveness of language model fine-tuning,” in Annual Meeting of the Association for Computational Linguistics , 2020
2020
Earlier work this paper cites.
L. Floridi and M. Chiriatti, “Gpt-3: Its nature, scope, limits, and consequences,” Minds and Machines , 2020
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of machine learning research , vol. 21, no. 140, pp. 1–67, 2020
2020
Earlier work this paper cites.
J. O. Zhang, A. Sax, A. Zamir, L. Guibas, and J. Malik, “Side-tuning: a baseline for network adaptation via additive side networks,” in European Conference on Computer Vision , 2020
2020
Earlier work this paper cites.
M. Ding, Z. Yang, W. Hong, W. Zheng, C. Zhou, D. Yin, J. Lin, X. Zou, Z. Shao, H. Yang et al. , “Cogview: Mastering text-to-image generation via transformers,” Advances in neural information processing systems , vol. 34, pp. 19 822–19 835, 2021
2021
Earlier work this paper cites.
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” in International conference on machine learning . PMLR, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
J. Gurrola-Ramos, O. Dalmau, and T. E. Alarcón, “A residual dense u-net neural network for image denoising,” IEEE Access , vol. 9, pp. 31 742–31 754, 2021
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Y. Gu, X. Han, Z. Liu, and M. Huang, “Ppt: Pre-trained prompt tuning for few-shot learning,” in Annual Meeting of the Association for Computational Linguistics , 2021
2021
Earlier work this paper cites.
Y. Mao, L. Mathias, R. Hou, A. Almahairi, H. Ma, J. Han, W. tau Yih, and M. Khabsa, “Unipelt: A unified framework for parameter-efficient language model tuning,” in Annual Meeting of the Association for Computational Linguistics , 2021
2021
Earlier work this paper cites.
Y.-L. Sung, V. Nair, and C. Raffel, “Training neural networks with fixed sparse masks,” in Neural Information Processing Systems , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Ansell, E. Ponti, A. Korhonen, and I. Vulic, “Composable sparse fine-tuning for cross-lingual transfer,” in Annual Meeting of the Association for Computational Linguistics , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
R. L. L. IV, I. Balavzevi’c, E. Wallace, F. Petroni, S. Singh, and S. Riedel, “Cutting down on prompts and parameters: Simple few-shot learning with language models,” in Findings , 2021
2021
Earlier work this paper cites.
T. Vu, B. Lester, N. Constant, R. Al-Rfou, and D. M. Cer, “Spot: Better frozen model adaptation through soft prompt transfer,” in Annual Meeting of the Association for Computational Linguistics , 2021
2021
Earlier work this paper cites.
P. Liu, Z.-F. Gao, W. X. Zhao, Z. Y. Xie, Z.-Y. Lu, and J. rong Wen, “Enabling lightweight fine-tuning for pre-trained language model compression based on matrix product operators,” in Annual Meeting of the Association for Computational Linguistics , 2021
2021
Earlier work this paper cites.
J. Davison, “Compacter: Efficient low-rank hypercomplex adapter layers,” in Neural Information Processing Systems , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Y.-L. Sung, J. Cho, and M. Bansal, “Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks,” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 5217–5227, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Y. Su, X. Wang, Y. Qin, C.-M. Chan, Y. Lin, H. Wang, K. Wen, Z. Liu, P. Li, J. Li, L. Hou, M. Sun, and J. Zhou, “On transferability of prompt tuning for natural language processing,” in North American Chapter of the Association for Computational Linguistics , 2021
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in International conference on machine learning . PMLR, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Y. Zhang, K. Zhou, and Z. Liu, “Neural prompt search,” arXiv preprint arXiv:2206.04673 , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
S. Chen, C. Ge, Z. Tong, J. Wang, Y. Song, J. Wang, and P. Luo, “Adaptformer: Adapting vision transformers for scalable visual recognition,” Advances in Neural Information Processing Systems , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
M. Jia, L. Tang, B.-C. Chen, C. Cardie, S. Belongie, B. Hariharan, and S.-N. Lim, “Visual prompt tuning,” in European Conference on Computer Vision , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
K. Zhou, J. Yang, C. C. Loy, and Z. Liu, “Learning to prompt for vision-language models,” International Journal of Computer Vision , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
M. Parovic, G. Glavas, I. Vulic, and A. Korhonen, “Bad-x: Bilingual adapters improve zero-shot cross-lingual transfer,” in North American Chapter of the Association for Computational Linguistics , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
H. Yuan, Z. Yuan, R. Gan, J. Zhang, Y. Xie, and S. Yu, “Biobart: Pretraining and evaluation of a biomedical generative language model,” in Workshop on Biomedical Natural Language Processing , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Cited alongside, same era.
B. Wang and A. Komatsuzaki, “Gpt-j-6b: a 6 billion parameter autoregressive language model (2021),” URL https://github.com/kingoflolz/mesh-transformer-jax , 2022
2022
Cited alongside, same era.
J. Pan, Z. Lin, X. Zhu, J. Shao, and H. Li, “St-adapter: Parameter-efficient image-to-video transfer learning,” Advances in Neural Information Processing Systems , 2022
2022
Cited alongside, same era.
Y.-C. Liu, C.-Y. Ma, J. Tian, Z. He, and Z. Kira, “Polyhistor: Parameter-efficient multi-task adaptation for dense vision tasks,” Advances in Neural Information Processing Systems , 2022
2022
Cited alongside, same era.
2023
Later among the works it cites.
W. Dong, S. Xue, X. Duan, and S. Han, “Prompt tuning inversion for text-driven image editing using diffusion models,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Sun, D. Fu, Y. Hu, S. Wang, R. Rassin, D.-C. Juan, D. Alon, C. Herrmann, S. van Steenkiste, R. Krishna et al. , “Dreamsync: Aligning text-to-image generation with image understanding feedback,” in Synthetic Data for Computer Vision Workshop@ CVPR 2024 , 2023
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
Z. Wang, X. Yu, Y. Rao, J. Zhou, and J. Lu, “P2p: Tuning pre-trained image models for point cloud analysis with point-to-pixel prompting,” Advances in neural information processing systems , 2022
2022
Cited alongside, same era.
Z. Bu, Y.-X. Wang, S. Zha, and G. Karypis, “Differentially private bias-term only fine-tuning of foundation models,” Advances in Neural Information Processing Systems , 2022
2022
Cited alongside, same era.
D. Lian, D. Zhou, J. Feng, and X. Wang, “Scaling & shifting your features: A new baseline for efficient model tuning,” Advances in Neural Information Processing Systems , 2022
2022
Cited alongside, same era.
E. Grassucci, A. Zhang, and D. Comminiello, “Phnns: Lightweight neural networks via parameterized hypercomplex convolutions,” IEEE Transactions on Neural Networks and Learning Systems , 2022
2022
Cited alongside, same era.
Z. Jiang, T. Chen, X. Chen, Y. Cheng, L. Zhou, L. Yuan, A. Awadallah, and Z. Wang, “Dna: Improving few-shot transfer learning with low-rank decomposition and alignment,” in European Conference on Computer Vision , 2022
2022
Cited alongside, same era.
K. Zhou, J. Yang, C. C. Loy, and Z. Liu, “Conditional prompt learning for vision-language models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022
2022
Cited alongside, same era.
M. Shu, W. Nie, D.-A. Huang, Z. Yu, T. Goldstein, A. Anandkumar, and C. Xiao, “Test-time prompt tuning for zero-shot generalization in vision-language models,” Advances in Neural Information Processing Systems , 2022
2022
Cited alongside, same era.
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Yu, J. Cho, P. Yadav, and M. Bansal, “Self-chained image-language model for video localization and question answering,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
J. Xu, X. Liu, Y. Wu, Y. Tong, Q. Li, M. Ding, J. Tang, and Y. Dong, “Imagereward: Learning and evaluating human preferences for text-to-image generation,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
H. Liu, C. Li, Y. Li, and Y. J. Lee, “Improved baselines with visual instruction tuning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 26 296–26 306
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Z. Tang, Z. Yang, C. Zhu, M. Zeng, and M. Bansal, “Any-to-any generation via composable diffusion,” Advances in Neural Information Processing Systems , 2024
2024
Later among the works it cites.
Z. Tang, Z. Yang, M. Khademi, Y. Liu, C. Zhu, and M. Bansal, “Codi-2: In-context interleaved and interactive any-to-any generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Gu, X. Wang, J. Z. Wu, Y. Shi, Y. Chen, Z. Fan, W. Xiao, R. Zhao, S. Chang, W. Wu et al. , “Mix-of-show: Decentralized low-rank adaptation for multi-concept customization of diffusion models,” Advances in Neural Information Processing Systems , 2024
2024
Later among the works it cites.
Q. Yu, J. He, X. Deng, X. Shen, and L.-C. Chen, “Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip,” Advances in Neural Information Processing Systems , 2024
2024
Later among the works it cites.
M. Fu, K. Zhu, and J. Wu, “Dtl: Disentangled transfer learning for visual recognition,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2024
2024
Later among the works it cites.
C. Mou, X. Wang, L. Xie, Y. Wu, J. Zhang, Z. Qi, and Y. Shan, “T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2024
2024
Later among the works it cites.
X. Guo, M. Zheng, L. Hou, Y. Gao, Y. Deng, P. Wan, D. Zhang, Y. Liu, W. Hu, Z. Zha et al. , “I2v-adapter: A general image-to-video adapter for diffusion models,” in ACM SIGGRAPH 2024 Conference Papers , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
P. Gao, S. Geng, R. Zhang, T. Ma, R. Fang, Y. Zhang, H. Li, and Y. Qiao, “Clip-adapter: Better vision-language models with feature adapters,” International Journal of Computer Vision , 2024
2024
Later among the works it cites.
H. Wang, J. Chang, Y. Zhai, X. Luo, J. Sun, Z. Lin, and Q. Tian, “Lion: Implicit vision prompt tuning,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Ren, Y. Zhou, J. Yang, J. Shi, D. Liu, F. Liu, M. Kwon, and A. Shrivastava, “Customize-a-video: One-shot motion customization of text-to-video diffusion models,” ECCV , 2024
2024
Later among the works it cites.
L. Song, Y. Chen, S. Yang, X. Ding, Y. Ge, Y.-C. Chen, and Y. Shan, “Low-rank approximation for sparse attention in multi-modal llms,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
C. Liu, K. Sun, Q. Zhou, Y. Duan, J. Shu, H. Kan, Z. Gu, and J. Hu, “Cpmi-chatglm: Parameter-efficient fine-tuning chatglm with chinese patent medicine instructions,” Scientific Reports , vol. 14, no. 1, p. 6403, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
H. H. Zhao, P. Wang, Y. Zhao, H. Luo, F. Wang, and M. Z. Shou, “Sct: A simple baseline for parameter-efficient fine-tuning via salient channels,” International Journal of Computer Vision , 2024
2024
Later among the works it cites.
Y. Xin, J. Du, Q. Wang, Z. Lin, and K. Yan, “Vmt-adapter: Parameter-efficient transfer learning for multi-task dense scene understanding,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2024
2024
Later among the works it cites.
S. Basu, S. Hu, D. Massiceti, and S. Feizi, “Strong baselines for parameter-efficient few-shot fine-tuning,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
S. Zhang, J. Zhang, H. Zhang, and L. Zhuo, “Rastformer: region-aware spatiotemporal transformer for visual homogenization recognition in short videos,” Neural Computing and Applications , 2024
2024
Later among the works it cites.
Y. Yao, A. Zhang, Z. Zhang, Z. Liu, T.-S. Chua, and M. Sun, “Cpt: Colorful prompt tuning for pre-trained vision-language models,” AI Open , 2024
2024
Later among the works it cites.
J. Bai, K. Gao, S. Min, S.-T. Xia, Z. Li, and W. Liu, “Badclip: Trigger-aware prompt learning for backdoor attacks on clip,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
G. Kim, S. Kim, and S. Lee, “Aapl: Adding attributes to prompt learning for vision-language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024
2024
Later among the works it cites.
K. Goswami, S. Karanam, P. Udhayanan, K. Joseph, and B. V. Srinivasan, “Copl: Contextual prompt learning for vision-language understanding,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2024
2024
Later among the works it cites.
Z. Xiao, J. Shen, M. M. Derakhshani, S. Liao, and C. G. Snoek, “Any-shift prompting for generalization over distributions,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024
2024
Later among the works it cites.
M. Dorkenwald, N. Barazani, C. G. Snoek, and Y. M. Asano, “Pin: Positional insert unlocks object localisation abilities in vlms,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024
2024
Later among the works it cites.
J. Silva-Rodriguez, S. Hajimiri, I. Ben Ayed, and J. Dolz, “A closer look at the few-shot adaptation of large vision-language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024
2024
Later among the works it cites.
H. Yao, R. Zhang, and C. Xu, “Tcp: Textual-based class-aware prompt tuning for visual-language model,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024
2024
Later among the works it cites.
J. Zhang, S. Wu, L. Gao, H. T. Shen, and J. Song, “Dept: Decoupled prompt tuning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024
2024
Later among the works it cites.
H. Shi, S. D. Dao, and J. Cai, “Llmformer: Large language model for open-vocabulary semantic segmentation,” International Journal of Computer Vision , 2024
2024
Later among the works it cites.
M. Wang, J. Xing, B. Jiang, J. Chen, J. Mei, X. Zuo, G. Dai, J. Wang, and Y. Liu, “A multimodal, multi-task adapting framework for video action recognition,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
X. Jin, B. Zhang, W. Gong, K. Xu, X. Deng, P. Wang, Z. Zhang, X. Shen, and J. Feng, “Mv-adapter: Multimodal video transfer learning for video text retrieval,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024
2024
Later among the works it cites.
X. Huang, Z. Huang, S. Li, W. Qu, T. He, Y. Hou, Y. Zuo, and W. Ouyang, “Frozen clip transformer is an efficient point cloud encoder,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2024
2024
Later among the works it cites.
X. Zhou, D. Liang, W. Xu, X. Zhu, Y. Xu, Z. Zou, and X. Bai, “Dynamic adapter meets prompt tuning: Parameter-efficient transfer learning for point cloud analysis,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024
2024
Later among the works it cites.
Y. Sun, J. Chen, S. Zhang, X. Zhang, Q. Chen, G. Zhang, E. Ding, J. Wang, and Z. Li, “Vrp-sam: Sam with visual reference prompt,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024
2024
Later among the works it cites.
S. Zhao, D. Chen, Y.-C. Chen, J. Bao, S. Hao, L. Yuan, and K.-Y. K. Wong, “Uni-controlnet: All-in-one control to text-to-image diffusion models,” Advances in Neural Information Processing Systems , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
L. Ran, X. Cun, J.-W. Liu, R. Zhao, S. Zijie, X. Wang, J. Keppo, and M. Z. Shou, “X-adapter: Adding universal compatibility of plugins for upgraded diffusion model,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
D. A. Hudson, D. Zoran, M. Malinowski, A. K. Lampinen, A. Jaegle, J. L. McClelland, L. Matthey, F. Hill, and A. Lerchner, “Soda: Bottleneck diffusion models for representation learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024
2024
Later among the works it cites.
Z. Xing, Q. Dai, H. Hu, Z. Wu, and Y.-G. Jiang, “Simda: Simple diffusion adapter for efficient video generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
R. Feng, W. Weng, Y. Wang, Y. Yuan, J. Bao, C. Luo, Z. Chen, and B. Guo, “Ccedit: Creative and controllable video editing via diffusion models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024
2024
Later among the works it cites.
K. Zhang, Y. Zhou, X. Xu, B. Dai, and X. Pan, “Diffmorpher: Unleashing the capability of diffusion models for image morphing,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
J. Guo, X. Xu, Y. Pu, Z. Ni, C. Wang, M. Vasu, S. Song, G. Huang, and H. Shi, “Smooth diffusion: Crafting smooth latent spaces in diffusion models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024
2024
Later among the works it cites.
M. Cai, H. Liu, S. K. Mustikovela, G. P. Meyer, Y. Chai, D. Park, and Y. J. Lee, “Vip-llava: Making large multimodal models understand arbitrary visual prompts,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024
2024
Later among the works it cites.
Z. Li, B. Yang, Q. Liu, Z. Ma, S. Zhang, J. Yang, Y. Sun, Y. Liu, and X. Bai, “Monkey: Image resolution and text label are important things for large multi-modal models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024
2024
Later among the works it cites.
Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang, “Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face,” Advances in Neural Information Processing Systems , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.