Fetching the paper…
Reading the bibliography…
Zero-shot Referring Image Segmentation (RIS) identifies the instance mask that best aligns with a specified referring expression without training and fine-tuning, significantly reducing the labor-intensive annotation process.
Referitgame: Referring to objects in photographs of natural scenes
Kazemzadeh, S.; Ordonez, V.; Matten, M.; and Berg, T. 2014 · 2014
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Mao, J.; Huang, J.; Toshev, A.; Camburu, O.; Yuille, A. L.; and Murphy, K. 2016 · 2016
Earlier work this paper cites.
Modeling context between objects for referring expression understanding
Nagaraja, V. K.; Morariu, V. I.; and Davis, L. S. 2016 · 2016
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; and Batra, D. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Earlier work this paper cites.
Clevr-ref+: Diagnosing visual reasoning with referring expressions
Liu, R.; Liu, C.; Bai, Y.; and Yuille, A. L. 2019 · 2019
Earlier work this paper cites.
Phrasecut: Language-based image segmentation in the wild
Wu, C.; Lin, Z.; Cohen, S.; Bui, T.; and Maji, S. 2020 · 2020
Earlier work this paper cites.
Vision-language transformer and query generation for referring segmentation
Ding, H.; Liu, C.; Wang, S.; and Jiang, X. 2021 · 2021
Earlier work this paper cites.
Locate then segment: A strong pipeline for referring image segmentation
Jing, Y.; Kong, T.; Wang, W.; Wang, L.; Li, L.; and Tan, T. 2021 · 2021
Earlier work this paper cites.
Align before fuse: Vision and language representation learning with momentum distillation
Li, J.; Selvaraju, R.; Gotmare, A.; Joty, S.; Xiong, C.; and Hoi, S. C. H. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
Maskgit: Masked generative image transformer
Chang, H.; Zhang, H.; Jiang, L.; Liu, C.; and Freeman, W. T. 2022 · 2022
Earlier work this paper cites.
Masked-attention mask transformer for universal image segmentation
Cheng, B.; Misra, I.; Schwing, A. G.; Kirillov, A.; and Girdhar, R. 2022 · 2022
Earlier work this paper cites.
Decoupling zero-shot semantic segmentation
Ding, J.; Xue, N.; Xia, G.-S.; and Dai, D. 2022 · 2022
Earlier work this paper cites.
VLMAE: Vision-language masked autoencoder
He, S.; Guo, T.; Dai, T.; Qiao, R.; Wu, C.; Shu, X.; and Ren, B. 2022 · 2022
Cited alongside, same era.
Restr: Convolution-free referring image segmentation using transformers
Kim, N.; Kim, D.; Lan, C.; Zeng, W.; and Kwak, S. 2022 · 2022
Cited alongside, same era.
mplug: Effective and efficient vision-language learning by cross-modal skip-connections
Li, C.; Xu, H.; Tian, J.; Wang, W.; Yan, M.; Bi, B.; Ye, J.; Chen, H.; Xu, G.; Cao, Z.; et al. 2022 · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Cited alongside, same era.
Reco: Retrieve and co-segment for zero-shot transfer
Shin, G.; Xie, W.; and Albanie, S. 2022 · 2022
Cited alongside, same era.
Open-vocabulary semantic segmentation with mask-adapted clip
Liang, F.; Wu, B.; Dai, X.; Li, K.; Zhao, Y.; Zhang, H.; Zhang, P.; Vajda, P.; and Marculescu, D. 2023 · 2023
Later among the works it cites.
Gres: Generalized referring expression segmentation
Liu, C.; Ding, H.; and Jiang, X. 2023 · 2023
Later among the works it cites.
Ref-diff: Zero-shot referring image segmentation with generative models
Ni, M.; Zhang, Y.; Feng, K.; Li, X.; Guo, Y.; and Zuo, W. 2023 · 2023
Later among the works it cites.
Text augmented spatial-aware zero-shot referring image segmentation
Suo, Y.; Zhu, L.; and Yang, Y. 2023 · 2023
Later among the works it cites.
Efficient Remote Sensing Transformer for Coastline Detection with Sentinel-2 Satellite Imagery
Wang, Y.; Zhao, R.; and Sun, Z. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Strudel, R.; Laptev, I.; and Schmid, C. 2022 · 2022
Cited alongside, same era.
Groupvit: Semantic segmentation emerges from text supervision
Xu, J.; De Mello, S.; Liu, S.; Byeon, W.; Breuel, T.; Kautz, J.; and Wang, X. 2022 · 2022
Cited alongside, same era.
Lavt: Language-aware vision transformer for referring image segmentation
Yang, Z.; Wang, J.; Tang, Y.; Chen, K.; Zhao, H.; and Torr, P. H. 2022 · 2022
Cited alongside, same era.
Coca: Contrastive captioners are image-text foundation models
Yu, J.; Wang, Z.; Vasudevan, V.; Yeung, L.; Seyedhosseini, M.; and Wu, Y. 2022 · 2022
Cited alongside, same era.
Extract free dense labels from clip
Zhou, C.; Loy, C. C.; and Dai, B. 2022 · 2022
Cited alongside, same era.
MeViS: A large-scale benchmark for video segmentation with motion expressions
Ding, H.; Liu, C.; He, S.; Jiang, X.; and Loy, C. C. 2023 · 2023
Cited alongside, same era.
Global knowledge calibration for fast open-vocabulary segmentation
Han, K.; Liu, Y.; Liew, J. H.; Ding, H.; Liu, J.; Wang, Y.; Tang, Y.; Yang, Y.; Feng, J.; Zhao, Y.; et al. 2023 · 2023
Cited alongside, same era.
Yu, S.; Seo, P. H.; and Son, J. 2023 · 2023
Later among the works it cites.
Self-calibrated clip for training-free open-vocabulary segmentation
Bai, S.; Liu, Y.; Han, Y.; Zhang, H.; and Tang, Y. 2024 · 2024
Later among the works it cites.
Zero-shot referring expression comprehension via structural similarity between images and captions
Han, Z.; Zhu, F.; Lao, Q.; and Jiang, H. 2024 · 2024
Later among the works it cites.
Lisa: Reasoning segmentation via large language model
Lai, X.; Tian, Z.; Chen, Y.; Li, Y.; Yuan, Y.; Liu, S.; and Jia, J. 2024 · 2024
Later among the works it cites.
LQMFormer: Language-aware Query Mask Transformer for Referring Image Segmentation
Shah, N. A.; VS, V.; and Patel, V. M. 2024 · 2024
Later among the works it cites.
GroundVLP: Harnessing Zero-Shot Visual Grounding from Vision-Language Pre-training and Open-Vocabulary Object Detection
Shen, H.; Zhao, T.; Zhu, M.; and Yin, J. 2024 · 2024
Later among the works it cites.
Clip as rnn: Segment countless visual concepts without training endeavor
Sun, S.; Li, R.; Torr, P.; Gu, X.; and Li, S. 2024 · 2024
Later among the works it cites.
Convolution Meets Transformer: Efficient Hybrid Transformer for Semantic Segmentation with Very High Resolution Imagery
Wang, Y.; Zhao, R.; Wei, S.; Ni, J.; Wu, M.; Luo, Y.; and Luo, C. 2024b · 2024
Later among the works it cites.
Language-aware vision transformer for referring segmentation
Yang, Z.; Wang, J.; Ye, X.; Tang, Y.; Chen, K.; Zhao, H.; and Torr, P. H. 2024 · 2024
Later among the works it cites.
A survey on open-vocabulary detection and segmentation: Past, present, and future
Zhu, C.; and Chen, L. 2024 · 2024
Later among the works it cites.