Fetching the paper…
Reading the bibliography…
In the domain of computer vision, Parameter-Efficient Tuning (PET) is increasingly replacing the traditional paradigm of pre-training followed by full fine-tuning.
Parameter-efficient transfer learning with diff pruning
Guo, D.; Rush, A. M.; and Kim, Y. 2020 · 2012
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
Kazemzadeh, S.; Ordonez, V.; Matten, M.; and Berg, T. 2014 · 2014
Earlier work this paper cites.
Modeling context in referring expressions
Yu, L.; Poirson, P.; Yang, S.; Berg, A. C.; and Berg, T. L. 2016 · 2016
Earlier work this paper cites.
Recurrent multimodal interaction for referring image segmentation
Liu, C.; Lin, Z.; Shen, X.; Yang, J.; Lu, X.; and Yuille, A. 2017 · 2017
Earlier work this paper cites.
Referring image segmentation via recurrent refinement networks
Li, R.; Li, K.; Kuo, Y.-C.; Shu, M.; Qi, X.; Shen, X.; and Jia, J. 2018 · 2018
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; De Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019 · 2019
Earlier work this paper cites.
Clip-adapter: Better vision-language models with feature adapters
Gao, P.; Geng, S.; Zhang, R.; Ma, T.; Fang, R.; Zhang, Y.; Li, H.; and Qiao, Y. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Earlier work this paper cites.
MDETR-modulated detection for end-to-end multi-modal understanding
Kamath, A.; Singh, M.; LeCun, Y.; Synnaeve, G.; Misra, I.; and Carion, N. 2021 · 2021
Earlier work this paper cites.
Compacter: Efficient low-rank hypercomplex adapter layers
Karimi Mahabadi, R.; Henderson, J.; and Ruder, S. 2021 · 2021
Earlier work this paper cites.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L.; and Liang, P. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
CRIS: CLIP-Driven Referring Image Segmentation
Wang, Z.; Lu, Y.; Li, Q.; Tao, X.; Guo, Y.; Gong, M.; and Liu, T. 2021 · 2021
Earlier work this paper cites.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Zaken, E. B.; Ravfogel, S.; and Goldberg, Y. 2021 · 2021
Earlier work this paper cites.
VLT: Vision-Language Transformer and Query Generation for Referring Segmentation
Ding, H.; Liu, C.; Wang, S.; and Jiang, X. 2022 · 2022
Cited alongside, same era.
ReSTR: Convolution-free Referring Image Segmentation Using Transformers
Kim, N.; Kim, D.; Lan, C.; Zeng, W.; and Kwak, S. 2022 · 2022
Cited alongside, same era.
Cris: Clip-driven referring image segmentation
Wang, Z.; Lu, Y.; Li, Q.; Tao, X.; Guo, Y.; Gong, M.; and Liu, T. 2022 · 2022
Cited alongside, same era.
Lavt: Language-aware vision transformer for referring image segmentation
Yang, Z.; Wang, J.; Tang, Y.; Chen, K.; Zhao, H.; and Torr, P. H. 2022 · 2022
Cited alongside, same era.
Learning to prompt for vision-language models
Zhou, K.; Yang, J.; Loy, C. C.; and Liu, Z. 2022 · 2022
Cited alongside, same era.
Contrastive grouping with transformer for referring image segmentation
Tang, J.; Zheng, G.; Shi, C.; and Yang, S. 2023 · 2023
Later among the works it cites.
BarLeRIa: An Efficient Tuning Framework for Referring Image Segmentation
Wang, Y.; Li, J.; ZHANG, X.; Shi, B.; Li, C.; Dai, W.; Xiong, H.; and Tian, Q. 2023 · 2023
Later among the works it cites.
Bridging vision and language encoders: Parameter-efficient tuning for referring image segmentation
Xu, Z.; Chen, Z.; Zhang, Y.; Song, Y.; Wan, X.; and Li, G. 2023 · 2023
Later among the works it cites.
Universal instance perception as object discovery and retrieval
Yan, B.; Jiang, Y.; Wu, J.; Wang, D.; Luo, P.; Yuan, Z.; and Lu, H. 2023 · 2023
Later among the works it cites.
Unleashing Text-to-Image Diffusion Models for Visual Perception
Zhao, W.; Rao, Y.; Liu, Z.; Liu, B.; Zhou, J.; and Lu, J. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chng, Y. X.; Zheng, H.; Han, Y.; Qiu, X.; and Huang, G. 2023 · 2023
Cited alongside, same era.
Eva: Exploring the limits of masked visual representation learning at scale
Fang, Y.; Wang, W.; Xie, B.; Sun, Q.; Wu, L.; Wang, X.; Huang, T.; Wang, X.; and Cao, Y. 2023 · 2023
Cited alongside, same era.
Consolidator: Mergable Adapter with Group Connections for Vision Transformer
Hao, T.; Chen, H.; Guo, Y.; and Ding, G. 2023 · 2023
Cited alongside, same era.
Camouflaged object detection with feature decomposition and edge reconstruction
He, C.; Li, K.; Zhang, Y.; Tang, L.; Zhang, Y.; Guo, Z.; and Li, X. 2023 · 2023
Cited alongside, same era.
Beyond One-to-One: Rethinking the Referring Image Segmentation
Hu, Y.; Wang, Q.; Shao, W.; Xie, E.; Li, Z.; Han, J.; and Luo, P. 2023 · 2023
Cited alongside, same era.
Lisa: Reasoning segmentation via large language model
Lai, X.; Tian, Z.; Chen, Y.; Li, Y.; Yuan, Y.; Liu, S.; and Jia, J. 2023 · 2023
Cited alongside, same era.
G2L: Semantically Aligned and Uniform Video Grounding via Geodesic and Game Theory
Li, H.; Cao, M.; Cheng, X.; Li, Y.; Zhu, Z.; and Zou, Y. 2023 · 2023
Cited alongside, same era.
Zou, X.; Yang, J.; Zhang, H.; Li, F.; Li, L.; Gao, J.; and Lee, Y. J. 2023 · 2023
Later among the works it cites.
Real-world Image Dehazing with Coherence-based Label Generator and Cooperative Unfolding Network
Fang, C.; He, C.; Xiao, F.; Zhang, Y.; Tang, L.; Zhang, Y.; Li, K.; and Li, X. 2024 · 2024
Later among the works it cites.
A survey of methods for addressing the challenges of referring image segmentation
Ji, L.; Du, Y.; Dang, Y.; Gao, W.; and Zhang, H. 2024 · 2024
Later among the works it cites.
Extending CLIP’s Image-Text Alignment to Referring Image Segmentation
Kim, S.; Kang, M.; Kim, D.; Park, J.; and Kwak, S. 2024 · 2024
Later among the works it cites.
DARA: Domain-and Relation-aware Adapters Make Parameter-efficient Tuning for Visual Grounding
Liu, T.; Liu, X.; Huang, S.; Chen, H.; Yin, Q.; Qin, L.; Wang, D.; and Hu, Y. 2024 · 2024
Later among the works it cites.
Uncertainty-aware sign language video retrieval with probability distribution modeling
Wu, X.; Li, H.; Luo, Y.; Cheng, X.; Zhuang, X.; Cao, M.; and Fu, K. 2024 · 2024
Later among the works it cites.
Remamber: Referring image segmentation with mamba twister
Yang, Y.; Ma, C.; Yao, J.; Zhong, Z.; Zhang, Y.; and Wang, Y. 2024 · 2024
Later among the works it cites.
Kdpror: A knowledge-decoupling probabilistic framework for video-text retrieval
Zhuang, X.; Li, H.; Cheng, X.; Zhu, Z.; Xie, Y.; and Zou, Y. 2025 · 2025
Closest in time.