Fetching the paper…
Reading the bibliography…
Open-vocabulary image segmentation aims to partition an image into semantic regions according to arbitrary text descriptions.
The hungarian method for the assignment problem
H. W. Kuhn · 1955
Earlier work this paper cites.
On seeing stuff: the perception of materials by humans and machines
E. H. Adelson · 2001
Earlier work this paper cites.
Computer vision: a modern approach
D. A. Forsyth and J. Ponce · 2002
Earlier work this paper cites.
Statistical region merging
R. Nock and F. Nielsen · 2004
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
The role of context for object detection and semantic segmentation in the wild
R. Mottaghi, X. Chen, X. Liu, N.-G. Cho, S.-W. Lee, S. Fidler, R. Urtasun, and A. Yuille · 2014
Earlier work this paper cites.
Image segmentation using k-means clustering algorithm and subtractive clustering algorithm
N. Dhanachandra, K. Manglem, and Y. J. Chanu · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
J. Long, E. Shelhamer, and T. Darrell · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Segmentation from natural language expressions
R. Hu, M. Rohrbach, and T. Darrell · 2016
Earlier work this paper cites.
Modeling context between objects for referring expression understanding
V. K. Nagaraja, V. I. Morariu, and L. S. Davis · 2016
Earlier work this paper cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Earlier work this paper cites.
Focal loss for dense object detection
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Generalised dice overlap as a deep learning loss function for highly unbalanced segmentations
C. H. Sudre, W. Li, T. Vercauteren, S. Ourselin, and M. Jorge Cardoso · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Mattnet: Modular attention network for referring expression comprehension
L. Yu, Z. Lin, X. Shen, J. Yang, X. Lu, M. Bansal, and T. L. Berg · 2018
Earlier work this paper cites.
Zero-shot semantic segmentation
M. Bucher, T.-H. Vu, M. Cord, and P. Pérez · 2019
Earlier work this paper cites.
Cascade r-cnn: high quality object detection and instance segmentation
Z. Cai and N. Vasconcelos · 2019
Earlier work this paper cites.
Panoptic segmentation
A. Kirillov, K. He, R. Girshick, C. Rother, and P. Dollár · 2019
Earlier work this paper cites.
Generalized intersection over union: A metric and a loss for bounding box regression
H. Rezatofighi, N. Tsoi, J. Gwak, A. Sadeghian, I. Reid, and S. Savarese · 2019
Earlier work this paper cites.
Objects365: A large-scale, high-quality dataset for object detection
S. Shao, Z. Li, T. Zhang, C. Peng, G. Yu, X. Zhang, J. Li, and J. Sun · 2019
Earlier work this paper cites.
Semantic projection network for zero-and few-label semantic segmentation
Y. Xian, S. Choudhury, Y. He, B. Schiele, and Z. Akata · 2019
Cited alongside, same era.
Semantic understanding of scenes through the ade20k dataset
B. Zhou, H. Zhao, X. Puig, T. Xiao, S. Fidler, A. Barriuso, and A. Torralba · 2019
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
Linguistic structure guided context modeling for referring image segmentation
T. Hui, S. Liu, S. Huang, G. Li, S. Yu, F. Zhang, and J. Han · 2020
Cited alongside, same era.
Deformable detr: Deformable transformers for end-to-end object detection
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai · 2020
Cited alongside, same era.
Mask dino: Towards a unified transformer-based framework for object detection and segmentation, 2022
F. Li, H. Zhang, H. xu, S. Liu, L. Zhang, L. M. Ni, and H.-Y. Shum · 2022
Later among the works it cites.
Grounded language-image pre-training
L. H. Li*, P. Zhang*, H. Zhang*, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J.-N. Hwang, K.-W. Chang, and J. Gao · 2022
Later among the works it cites.
Exploring plain vision transformer backbones for object detection
Y. Li, H. Mao, R. Girshick, and K. He · 2022
Later among the works it cites.
Open-vocabulary semantic segmentation with mask-adapted clip
F. Liang, B. Wu, X. Dai, K. Li, Y. Zhao, H. Zhang, P. Zhang, P. Vajda, and D. Marculescu · 2022
Later among the works it cites.
Denseclip: Language-guided dense prediction with context-aware prompting
Y. Rao, W. Zhao, G. Chen, Y. Tang, Z. Zhu, G. Huang, J. Zhou, and J. Lu · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Part-aware panoptic segmentation
D. de Geus, P. Meletis, C. Lu, X. Wen, and G. Dubbelman · 2021
Cited alongside, same era.
Vision-language transformer and query generation for referring segmentation
H. Ding, C. Liu, S. Wang, and X. Jiang · 2021
Cited alongside, same era.
Encoder fusion network with co-attention embedding for referring image segmentation
G. Feng, Z. Hu, L. Zhang, and H. Lu · 2021
Cited alongside, same era.
Yolox: Exceeding yolo series in 2021
Z. Ge, S. Liu, F. Wang, Z. Li, and J. Sun · 2021
Cited alongside, same era.
Locate then segment: A strong pipeline for referring image segmentation
Y. Jing, T. Kong, W. Wang, L. Wang, L. Li, and T. Tan · 2021
Cited alongside, same era.
Referring transformer: A one-step approach to multi-task visual grounding
M. Li and L. Sigal · 2021
Cited alongside, same era.
Referring transformer: A one-step approach to multi-task visual grounding
L. Muchen and S. Leonid · 2021
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Later among the works it cites.
Computer vision: algorithms and applications
R. Szeliski · 2022
Later among the works it cites.
Cris: Clip-driven referring image segmentation
Z. Wang, Y. Lu, Q. Li, X. Tao, Y. Guo, M. Gong, and T. Liu · 2022
Later among the works it cites.
Towards robust referring image segmentation
J. Wu, X. Li, X. Li, H. Ding, Y. Tong, and D. Tao · 2022
Later among the works it cites.
Groupvit: Semantic segmentation emerges from text supervision
J. Xu, S. De Mello, S. Liu, W. Byeon, T. Breuel, J. Kautz, and X. Wang · 2022
Later among the works it cites.
A simple baseline for open-vocabulary semantic segmentation with pre-trained vision-language model
M. Xu, Z. Zhang, F. Wei, Y. Lin, Y. Cao, H. Hu, and X. Bai · 2022
Later among the works it cites.
Lavt: Language-aware vision transformer for referring image segmentation
Z. Yang, J. Wang, Y. Tang, K. Chen, H. Zhao, and P. H. Torr · 2022
Later among the works it cites.
Dino: Detr with improved denoising anchor boxes for end-to-end object detection
H. Zhang, F. Li, S. Liu, L. Zhang, H. Su, J. Zhu, L. M. Ni, and H.-Y. Shum · 2022
Later among the works it cites.
Generalized decoding for pixel, image, and language
X. Zou, Z.-Y. Dou, J. Yang, Z. Gan, L. Li, C. Li, X. Dai, H. Behl, J. Wang, L. Yuan, et al · 2022
Later among the works it cites.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al · 2023
Closest in time.
Polyformer: Referring image segmentation as sequential polygon generation
J. Liu, H. Ding, Z. Cai, Y. Zhang, R. K. Satzoda, V. Mahadevan, and R. Manmatha · 2023
Closest in time.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, C. Li, J. Yang, H. Su, J. Zhu, et al · 2023
Closest in time.
Learning open-vocabulary semantic segmentation models from natural language supervision
J. Xu, J. Hou, Y. Zhang, R. Feng, Y. Wang, Y. Qiao, and W. Xie · 2023
Closest in time.
Open-vocabulary panoptic segmentation with text-to-image diffusion models
J. Xu, S. Liu, A. Vahdat, W. Byeon, X. Wang, and S. De Mello · 2023
Closest in time.
Universal instance perception as object discovery and retrieval
B. Yan, Y. Jiang, J. Wu, D. Wang, P. Luo, Z. Yuan, and H. Lu · 2023
Closest in time.
Unleashing text-to-image diffusion models for visual perception
W. Zhao, Y. Rao, Z. Liu, B. Liu, J. Zhou, and J. Lu · 2023
Closest in time.
Segment everything everywhere all at once
X. Zou, J. Yang, H. Zhang, F. Li, L. Li, J. Gao, and Y. J. Lee · 2023
Closest in time.