Fetching the paper…
Reading the bibliography…
Referring Image Segmentation (RIS) is a challenging task that requires an algorithm to segment objects referred by free-form language expressions.
Dimensionality reduction by learning an invariant mapping
R. Hadsell, S. Chopra, and Y. LeCun · 2006
Earlier work this paper cites.
Imagespirit: Verbal guided image parsing
M. Cheng, S. Zheng, W. Lin, V. Vineet, P. Sturgess, N. Crook, N. J. Mitra, and P. H. Torr · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
J. Long, E. Shelhamer, and T. Darrell · 2015
Earlier work this paper cites.
Instance-aware semantic segmentation via multi-task network cascades
J. Dai, K. He, and J. Sun · 2016
Earlier work this paper cites.
Segmentation from natural language expressions
R. Hu, M. Rohrbach, and T. Darrell · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. Yuille, and K. Murphy · 2016
Earlier work this paper cites.
Modeling context between objects for referring expression understanding
V. K. Nagaraja, V. I. Morariu, and L. S. Davis · 2016
Earlier work this paper cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Earlier work this paper cites.
Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs
L. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. Yuille · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
Mask r-cnn
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Earlier work this paper cites.
Recurrent multimodal interaction for referring image segmentation
C. Liu, Z. Lin, X. Shen, J. Yang, X. Lu, and A. Yuille · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Interactive text2pickup networks for natural language-based human–robot collaboration
H. Ahn, S. Choi, N. Kim, G. Cha, and S. Oh · 2018
Earlier work this paper cites.
Referring image segmentation via recurrent refinement networks
R. Li, K. Li, Y. Kuo, M. Shu, X. Qi, X. Shen, and J. Jia · 2018
Earlier work this paper cites.
Key-word-aware network for referring expression image segmentation
H. Shi, H. Li, F. Meng, and Q. Wu · 2018
Earlier work this paper cites.
Mattnet: Modular attention network for referring expression comprehension
L. Yu, Z. Lin, X. Shen, J. Yang, X. Lu, M. Bansal, and T. L. Berg · 2018
Earlier work this paper cites.
Yolact: Real-time instance segmentation
D. Bolya, C. Zhou, F. Xiao, and Y. J. Lee · 2019
Earlier work this paper cites.
See-through-text grouping for referring image segmentation
D. Chen, S. Jia, Y. Lo, H. Chen, and T. Liu · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Panoptic segmentation
A. Kirillov, K. He, R. Girshick, C. Rother, and P. Dollár · 2019
Earlier work this paper cites.
Cross-modal self-attention network for referring image segmentation
L. Ye, M. Rochan, Z. Liu, and Y. Wang · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2020
Cited alongside, same era.
Panoptic-deeplab: A simple, strong, and fast baseline for bottom-up panoptic segmentation
B. Cheng, M. Collins, Y. Zhu, T. Liu, T. Huang, A. Hartwig, and L. Chen · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Cited alongside, same era.
Contrastive learning for label efficient semantic segmentation
X. Zhao, R. Vemulapalli, P. A. Mansfield, B.Gong, B. Green, L. Shapira, and Y. Wu · 2021
Later among the works it cites.
Scaling instruction-finetuned language models
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, E. Li, X. Wang, M. Dehghani, S. Brahma, et al · 2022
Later among the works it cites.
Latency-aware spatial-wise dynamic networks
Y. Han, Z. Yuan, Y. Pu, C. Xue, S. Song, G. Sun, and G. Huang · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2022
Later among the works it cites.
Restr: Convolution-free referring image segmentation using transformers
N. Kim, D. Kim, C. Lan, W. Zeng, and S. Kwak · 2022
Later among the works it cites.
Grounded language-image pre-training
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bidirectional relationship inferring network for referring image localization and segmentation
Z. Hu, G. Feng, J. Sun, L. Zhang, and H. Lu · 2020
Cited alongside, same era.
Referring image segmentation via cross-modal progressive comprehension
S. Huang, T. Hui, S. Liu, G. Li, Y. Wei, J. Han, L. Liu, and B. Li · 2020
Cited alongside, same era.
Linguistic structure guided context modeling for referring image segmentation
T. Hui, S. Liu, S. Huang, G. Li, S. Yu, F. Zhang, and J. Han · 2020
Cited alongside, same era.
Vl-bert: Pre-training of generic visual-linguistic representations
W. Su, X. Zhu, Y. Cao, B. Li, L. Lu, F. Wei, and J. Dai · 2020
Cited alongside, same era.
Exploring simple siamese representation learning
X. Chen and K. He · 2021
Cited alongside, same era.
Vision-language transformer and query generation for referring segmentation
H. Ding, C. Liu, S. Wang, and X. Jiang · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Cited alongside, same era.
L. Li, P. Zhang, H. Zhang, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J. Hwang, K. Chang, and J. Gao · 2022
Later among the works it cites.
Cris: Clip-driven referring image segmentation
Z. Wang, Y. Lu, Q. Li, X. Tao, Y. Guo, M. Gong, and T. Liu · 2022
Later among the works it cites.
Lavt: Language-aware vision transformer for referring image segmentation
Z. Yang, J. Wang, Y. Tang, K. Chen, H. Zhao, and P. H. Torr · 2022
Later among the works it cites.
Seqtr: A simple yet universal network for visual grounding
C. Zhu, Y. Zhou, Y. Shen, G. Luo, X. Pan, M. Lin, C. Chen, L. Cao, X. Sun, and R. Ji · 2022
Later among the works it cites.
Beyond one-to-one: Rethinking the referring image segmentation
Y. Hu, Qi Wang, W. Shao, E. Xie, Z. Li, J. Han, and P. Luo · 2023
Closest in time.
Masked vision and language modeling for multi-modal representation learning
G. Kwon, Z. Cai, A. Ravichandran, E. Bas, R. Bhotika, and S. Soatto · 2023
Closest in time.
Lisa: Reasoning segmentation via large language model
Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia · 2023
Closest in time.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
J. Li, D. Li, S. Savarese, and S. Hoi · 2023
Closest in time.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Closest in time.
Image as a foreign language: Beit pretraining for all vision and vision-language tasks
W. Wang, H. Bao, L. Dong, J. Bjorck, Z. Peng, Q. Liu, K. Aggarwal, O. K. Mohammed, S. Singhal, S. Som, and F. Wei · 2023
Closest in time.
Meta compositional referring expression segmentation
L. Xu, M. H. Huang, X. Shang, Z. Yuan, Y. Sun, and J. Liu · 2023
Closest in time.
Mmnet: Multi-mask network for referring image segmentation
Y. Yan, X. He, W. Wan, and J. Liu · 2023
Closest in time.
Semantics-aware dynamic localization and refinement for referring image segmentation
Z. Yang, J. Wang, Y. Tang, K. Chen, H. Zhao, and P. H. Torr · 2023
Closest in time.
Unleashing text-to-image diffusion models for visual perception
W. Zhao, Y. Rao, Z. Liu, B. Liu, J. Zhou, and J. Lu · 2023
Closest in time.
Segment everything everywhere all at once
X. Zou, J. Yang, H. Zhang, F. Li, L. Li, J. Gao, and Y. Lee · 2023
Closest in time.
Parallel vertex diffusion for unified visual grounding
Z. Cheng, K. Li, P. Jin, X. Ji, L. Yuan, C. Liu, and J. Chen · 2024
Closest in time.
Revisiting non-autoregressive transformers for efficient image synthesis
Z. Ni, Y. Wang, R. Zhou, J. Guo, J. Hu, Z. Liu, S. Song, Y. Yao, and G. Huang · 2024
Closest in time.