Fetching the paper…
Reading the bibliography…
Visual segmentation seeks to partition images, video frames, or point clouds into multiple segments or groups.
H. W. Kuhn, “The hungarian method for the assignment problem,” Naval research logistics quarterly , 1955
1955
Earlier work this paper cites.
M. Kass, A. Witkin, and D. Terzopoulos, “Snakes: Active contour models,” IJCV , 1988
1988
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , 1997
1997
Earlier work this paper cites.
J. Shi and J. Malik, “Normalized cuts and image segmentation,” TPAMI , 2000
2000
Earlier work this paper cites.
J. Malik, S. Belongie, T. Leung, and J. Shi, “Contour and texture analysis for image segmentation,” IJCV , 2001
2001
Earlier work this paper cites.
X. Y. Stella and J. Shi, “Multiclass spectral clustering,” in ICCV , 2003
2003
Earlier work this paper cites.
F. Schroff, A. Criminisi, and A. Zisserman, “Object class segmentation using random forests.” in BMVC , 2008
2008
Earlier work this paper cites.
M. Everingham, L. J. V. Gool, C. K. I. Williams, J. M. Winn, and A. Zisserman, “The PASCAL visual object classes (voc) challenge,” IJCV , 2010
2010
Earlier work this paper cites.
V. Ordonez, G. Kulkarni, and T. Berg, “Im2text: Describing images using 1 million captioned photographs,” NeurIPS , 2011
2011
Earlier work this paper cites.
R. Mottaghi, X. Chen, X. Liu, N.-G. Cho, S.-W. Lee, S. Fidler, R. Urtasun, and A. Yuille, “The role of context for object detection and semantic segmentation in the wild,” in CVPR , 2014
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft COCO: Common objects in context,” in ECCV , 2014
2014
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “Imagenet large scale visual recognition challenge,” IJCV , 2015
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in ICLR , 2015
2015
Earlier work this paper cites.
J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in CVPR , 2015
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” in NeurIPS , 2015
2015
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional networks for biomedical image segmentation,” in MICCAI , 2015
2015
Earlier work this paper cites.
J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. Shamma, M. Bernstein, and L. Fei-Fei, “Image retrieval using scene graphs,” in CVPR , 2015
2015
Earlier work this paper cites.
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick, “Microsoft coco captions: Data collection and evaluation server,” CVPR , 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016
2016
Earlier work this paper cites.
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” in CVPR , 2016
2016
Earlier work this paper cites.
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling context in referring expressions,” in ECCV , 2016
2016
Earlier work this paper cites.
E. Shelhamer, K. Rakelly, J. Hoffman, and T. Darrell, “Clockwork convnets for video semantic segmentation,” in ECCV , 2016
2016
Earlier work this paper cites.
F. Perazzi, J. Pont-Tuset, B. McWilliams, L. Van Gool, M. Gross, and A. Sorkine-Hornung, “A benchmark dataset and evaluation methodology for video object segmentation,” in CVPR , 2016
2016
Earlier work this paper cites.
F. Milletari, N. Navab, and S. Ahmadi, “V-Net: Fully convolutional neural networks for volumetric medical image segmentation,” in 3DV , 2016
2016
Earlier work this paper cites.
L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs,” IEEE TPAMI , vol. 40, no. 4, pp. 834–848, 2017
2017
Earlier work this paper cites.
H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid scene parsing network,” in CVPR , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in NIPS , 2017
2017
Earlier work this paper cites.
B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba, “Semantic understanding of scenes through the ADE20K dataset,” CVPR , 2017
2017
Earlier work this paper cites.
G. Neuhold, T. Ollmann, S. Rota Bulo, and P. Kontschieder, “The mapillary vistas dataset for semantic understanding of street scenes,” in ICCV , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
C. Peng, X. Zhang, G. Yu, G. Luo, and J. Sun, “Large kernel matters–improve semantic segmentation by global convolutional network,” in CVPR , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask R-CNN,” in ICCV , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
R. Gadde, V. Jampani, and P. V. Gehler, “Semantic video cnns through representation warping,” in ICCV , 2017
2017
Earlier work this paper cites.
X. Zhu, Y. Xiong, J. Dai, L. Yuan, and Y. Wei, “Deep feature flow for video recognition,” in CVPR , 2017
2017
Earlier work this paper cites.
C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “PointNet: Deep learning on point sets for 3d classification and segmentation,” in CVPR , 2017, pp. 652–660
2017
Earlier work this paper cites.
C. R. Qi, L. Yi, H. Su, and L. J. Guibas, “Pointnet++: Deep hierarchical feature learning on point sets in a metric space,” NeurIPS , 2017
2017
Earlier work this paper cites.
T.-Y. Lin, P. Dollár, R. B. Girshick, K. He, B. Hariharan, and S. J. Belongie, “Feature pyramid networks for object detection,” in CVPR , 2017
2017
Earlier work this paper cites.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in ICCV , 2017
2017
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma et al. , “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” IJCV , 2017
2017
Earlier work this paper cites.
H. Ding, X. Jiang, B. Shuai, A. Q. Liu, and G. Wang, “Context contrasted feature and gated multi-scale aggregation for scene segmentation,” in CVPR , 2018
2018
Earlier work this paper cites.
H. Zhao, Y. Zhang, S. Liu, J. Shi, C. Change Loy, D. Lin, and J. Jia, “PSANet: Point-wise spatial attention network for scene parsing,” in ECCV , 2018
2018
Earlier work this paper cites.
X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in CVPR , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
C. Yu, J. Wang, C. Peng, C. Gao, G. Yu, and N. Sang, “Learning a discriminative feature network for semantic segmentation,” in CVPR , 2018
2018
Earlier work this paper cites.
B. Shuai, H. Ding, T. Liu, G. Wang, and X. Jiang, “Toward achieving robust low-level and high-level scene parsing,” IEEE TIP , vol. 28, no. 3, pp. 1378–1390, 2018
2018
Earlier work this paper cites.
C. Yu, J. Wang, C. Peng, C. Gao, G. Yu, and N. Sang, “BiSeNet: Bilateral segmentation network for real-time semantic segmentation,” in ECCV , 2018
2018
Earlier work this paper cites.
W. Wang, R. Yu, Q. Huang, and U. Neumann, “SGPN: Similarity group proposal network for 3d point cloud instance segmentation,” in CVPR , 2018
2018
Earlier work this paper cites.
Y. Chen, Y. Kalantidis, J. Li, S. Yan, and J. Feng, “A2-Nets: Double attention networks,” NeurIPS , 2018
2018
Earlier work this paper cites.
L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam, “Encoder-decoder with atrous separable convolution for semantic image segmentation,” in ECCV , 2018
2018
Earlier work this paper cites.
L. Landrieu and M. Simonovsky, “Large-scale point cloud semantic segmentation with superpoint graphs,” in CVPR , 2018
2018
Earlier work this paper cites.
H. Zhao, X. Qi, X. Shen, J. Shi, and J. Jia, “ICNet for real-time semantic segmentation on high-resolution images,” ECCV , 2018
2018
Earlier work this paper cites.
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in ACL) , 2018
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in NAACL , 2019
2019
Earlier work this paper cites.
H. Hu, Z. Zhang, Z. Xie, and S. Lin, “Local relation networks for image recognition,” in ICCV , 2019
2019
Earlier work this paper cites.
L. Yang, Y. Fan, and N. Xu, “Video instance segmentation,” in ICCV , 2019
2019
Earlier work this paper cites.
H. Ding, X. Jiang, B. Shuai, A. Q. Liu, and G. Wang, “Semantic correlation promoted shape-variant context for segmentation,” in CVPR , 2019
2019
Earlier work this paper cites.
L. Zhang, X. Li, A. Arnab, K. Yang, Y. Tong, and P. H. Torr, “Dual graph convolutional network for semantic segmentation,” in BMVC , 2019
2019
Earlier work this paper cites.
H. Ding, X. Jiang, A. Q. Liu, N. M. Thalmann, and G. Wang, “Boundary-aware feature propagation for scene segmentation,” in ICCV , 2019
2019
Earlier work this paper cites.
J. Fu, J. Liu, H. Tian, Z. Fang, and H. Lu, “Dual attention network for scene segmentation,” in CVPR , 2019
2019
Earlier work this paper cites.
D. Neven, B. D. Brabandere, M. Proesmans, and L. V. Gool, “Instance segmentation by jointly optimizing spatial embeddings and clustering bandwidth,” in CVPR , 2019
2019
Earlier work this paper cites.
K. Chen, J. Pang, J. Wang, Y. Xiong, X. Li, S. Sun, W. Feng, Z. Liu, J. Shi, W. Ouyang, C. C. Loy, and D. Lin, “Hybrid task cascade for instance segmentation,” in CVPR , 2019
2019
Earlier work this paper cites.
D. Bolya, C. Zhou, F. Xiao, and Y. J. Lee, “YOLACT: Real-time instance segmentation,” in ICCV , 2019
2019
Earlier work this paper cites.
X. Chen, R. Girshick, K. He, and P. Dollár, “Tensormask: A foundation for dense object segmentation,” in ICCV , 2019
2019
Earlier work this paper cites.
Y. Xiong, R. Liao, H. Zhao, R. Hu, M. Bai, E. Yumer, and R. Urtasun, “UPSNet: A unified panoptic segmentation network,” in CVPR , 2019
2019
Earlier work this paper cites.
X. Zhu, H. Hu, S. Lin, and J. Dai, “Deformable convnets v2: More deformable, better results,” in CVPR , 2019
2019
Earlier work this paper cites.
P. Voigtlaender, M. Krause, A. Osep, J. Luiten, B. B. G. Sekar, A. Geiger, and B. Leibe, “MOTS: Multi-object tracking and segmentation,” in CVPR , 2019
2019
Earlier work this paper cites.
B. Yang, J. Wang, R. Clark, Q. Hu, S. Wang, A. Markham, and N. Trigoni, “Learning object bounding boxes for 3d instance segmentation on point clouds,” in NeurIPS , 2019
2019
Earlier work this paper cites.
L. Yi, W. Zhao, H. Wang, M. Sung, and L. J. Guibas, “GSPN: Generative shape proposal network for 3d instance segmentation in point cloud,” in CVPR , 2019
2019
Earlier work this paper cites.
J. Mao, X. Wang, and H. Li, “Interpolated convolutional networks for 3d point cloud understanding,” in ICCV , 2019
2019
Earlier work this paper cites.
G. Ghiasi, T.-Y. Lin, and Q. V. Le, “NAS-FPN: Learning scalable feature pyramid architecture for object detection,” in CVPR , 2019
2019
Earlier work this paper cites.
Y. Wu, A. Kirillov, F. Massa, W.-Y. Lo, and R. Girshick, “Detectron2,” https://github.com/facebookresearch/detectron2 , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
C.-C. Hsu, K.-J. Hsu, C.-C. Tsai, Y.-Y. Lin, and Y.-Y. Chuang, “Weakly supervised instance segmentation using the bounding box tightness prior,” NeurIPS , 2019
2019
Earlier work this paper cites.
S. W. Oh, J.-Y. Lee, N. Xu, and S. J. Kim, “Video object segmentation using space-time memory networks,” in ICCV , 2019
2019
Earlier work this paper cites.
S. Shao, Z. Li, T. Zhang, C. Peng, G. Yu, X. Zhang, J. Li, and J. Sun, “Objects365: A large-scale, high-quality dataset for object detection,” in ICCV , 2019
2019
Earlier work this paper cites.
K. Chen, J. Wang, J. Pang, Y. Cao, Y. Xiong, X. Li, S. Sun, W. Feng, Z. Liu, J. Xu et al. , “Mmdetection: Open mmlab detection toolbox and benchmark,” arXiv preprint , 2019
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” in NeurIPS , 2020
2020
Earlier work this paper cites.
H. Zhao, J. Jia, and V. Koltun, “Exploring self-attention for image recognition,” in CVPR , 2020
2020
Earlier work this paper cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in ECCV , 2020
2020
Earlier work this paper cites.
S. Hao, Y. Zhou, and Y. Guo, “A brief survey on semantic segmentation with deep learning,” Neurocomputing , 2020
2020
Earlier work this paper cites.
D. Kim, S. Woo, J.-Y. Lee, and I. S. Kweon, “Video panoptic segmentation,” in CVPR , 2020
2020
Earlier work this paper cites.
H. Ding, X. Jiang, B. Shuai, A. Q. Liu, and G. Wang, “Semantic segmentation with context encoding and multi-path decoding,” TIP , 2020
2020
Earlier work this paper cites.
X. Li, H. Zhao, L. Han, Y. Tong, S. Tan, and K. Yang, “Gated fully fusion for semantic segmentation,” in AAAI , 2020
2020
Earlier work this paper cites.
Y. Yuan, X. Chen, and J. Wang, “Object-contextual representations for semantic segmentation,” ECCV , 2020
2020
Earlier work this paper cites.
X. Li, A. You, Z. Zhu, H. Zhao, M. Yang, K. Yang, and Y. Tong, “Semantic flow for fast and accurate scene parsing,” in ECCV , 2020
2020
Earlier work this paper cites.
A. Kirillov, Y. Wu, K. He, and R. Girshick, “Pointrend: Image segmentation as rendering,” in CVPR , 2020
2020
Earlier work this paper cites.
X. Li, X. Li, L. Zhang, G. Cheng, J. Shi, Z. Lin, S. Tan, and Y. Tong, “Improving semantic segmentation via decoupled body and edge supervision,” in ECCV , 2020
2020
Earlier work this paper cites.
Z. Tian, C. Shen, and H. Chen, “Conditional convolutions for instance segmentation,” in ECCV , 2020
2020
Earlier work this paper cites.
R. Zhang, Z. Tian, C. Shen, M. You, and Y. Yan, “Mask encoding for single shot instance segmentation,” in CVPR , 2020
2020
Earlier work this paper cites.
B. Cheng, M. D. Collins, Y. Zhu, T. Liu, T. S. Huang, H. Adam, and L.-C. Chen, “Panoptic-DeepLab: A simple, strong, and fast baseline for bottom-up panoptic segmentation,” in CVPR , 2020
2020
Earlier work this paper cites.
X. Wang, R. Zhang, T. Kong, L. Li, and C. Shen, “SOLOv2: Dynamic and fast instance segmentation,” in NeurIPS , 2020
2020
Earlier work this paper cites.
H. Wang, Y. Zhu, B. Green, H. Adam, A. Yuille, and L.-C. Chen, “Axial-DeepLab: Stand-alone axial-attention for panoptic segmentation,” in ECCV , 2020
2020
Earlier work this paper cites.
G. Bertasius and L. Torresani, “Classifying, segmenting, and tracking object instances in video with mask propagation,” in CVPR , 2020
2020
Earlier work this paper cites.
L. Jiang, H. Zhao, S. Shi, S. Liu, C.-W. Fu, and J. Jia, “PointGroup: Dual-set point grouping for 3d instance segmentation,” in CVPR , 2020
2020
Earlier work this paper cites.
Q. Hu, B. Yang, L. Xie, S. Rosa, Y. Guo, Z. Wang, N. Trigoni, and A. Markham, “RandLA-Net: Efficient semantic segmentation of large-scale point clouds,” in CVPR , 2020
2020
Earlier work this paper cites.
D. Zhang, H. Zhang, J. Tang, M. Wang, X. Hua, and Q. Sun, “Feature pyramid transformer,” in ECCV , 2020
2020
Earlier work this paper cites.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” ICML , 2020
2020
Earlier work this paper cites.
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in CVPR , 2020
2020
Earlier work this paper cites.
P. Sun, J. Cao, Y. Jiang, R. Zhang, E. Xie, Z. Yuan, C. Wang, and P. Luo, “TransTrack: Multiple-object tracking with transformer,” arXiv preprint arXiv: 2012.15460 , 2020
2020
Earlier work this paper cites.
G. Luo, Y. Zhou, X. Sun, L. Cao, C. Wu, C. Deng, and R. Ji, “Multi-task collaborative network for joint referring expression comprehension and segmentation,” in CVPR , 2020
2020
Earlier work this paper cites.
Z. Liu, Z. Miao, X. Pan, X. Zhan, D. Lin, S. X. Yu, and B. Gong, “Open compound domain adaptation,” in CVPR , 2020
2020
Earlier work this paper cites.
Y. Yang and S. Soatto, “FDA: Fourier domain adaptation for semantic segmentation,” in CVPR , 2020
2020
Earlier work this paper cites.
J. Lambert, Z. Liu, O. Sener, J. Hays, and V. Koltun, “MSeg: A composite dataset for multi-domain semantic segmentation,” in CVPR , 2020
2020
Earlier work this paper cites.
Y. Wang, J. Zhang, M. Kan, S. Shan, and X. Chen, “Self-supervised equivariant attention mechanism for weakly supervised semantic segmentation,” in CVPR , 2020
2020
Earlier work this paper cites.
H. K. Cheng, J. Chung, Y.-W. Tai, and C.-K. Tang, “CascadePSP: Toward class-agnostic and very high-resolution segmentation via global and local refinement,” in CVPR , 2020
2020
Earlier work this paper cites.
H. Ding, S. Cohen, B. Price, and X. Jiang, “Phraseclick: toward achieving flexible interactive segmentation by phrase and click,” in ECCV . Springer, 2020, pp. 417–435
2020
Earlier work this paper cites.
M. Contributors, “MMSegmentation: Openmmlab semantic segmentation toolbox and benchmark,” https://github.com/open-mmlab/mmsegmentation , 2020
2020
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR , 2021
2021
Earlier work this paper cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” ICCV , 2021
2021
Earlier work this paper cites.
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable DETR: Deformable transformers for end-to-end object detection,” in ICLR , 2021
2021
Earlier work this paper cites.
H. Wang, Y. Zhu, H. Adam, A. Yuille, and L.-C. Chen, “MaX-DeepLab: End-to-end panoptic segmentation with mask transformers,” in CVPR , 2021
2021
Earlier work this paper cites.
H. Chen, Y. Wang, T. Guo, C. Xu, Y. Deng, Z. Liu, S. Ma, C. Xu, C. Xu, and W. Gao, “Pre-trained image processing transformer,” in CVPR , 2021, pp. 12 299–12 310
2021
Earlier work this paper cites.
G. Bertasius, H. Wang, and L. Torresani, “Is space-time attention all you need for video understanding?” in ICML , 2021
2021
Earlier work this paper cites.
H. Zhao, L. Jiang, J. Jia, P. H. Torr, and V. Koltun, “Point transformer,” in ICCV , 2021
2021
Earlier work this paper cites.
X. Pan, Z. Xia, S. Song, L. E. Li, and G. Huang, “3D object detection with pointformer,” in CVPR , 2021, pp. 7463–7472
2021
Earlier work this paper cites.
S. Minaee, Y. Y. Boykov, F. Porikli, A. J. Plaza, N. Kehtarnavaz, and D. Terzopoulos, “Image segmentation using deep learning: A survey,” TPAMI , 2021
2021
Earlier work this paper cites.
J. Miao, Y. Wei, Y. Wu, C. Liang, G. Li, and Y. Yang, “VSPW: A large-scale dataset for video scene parsing in the wild,” in CVPR , 2021
2021
Earlier work this paper cites.
M. Weber, J. Xie, M. Collins, Y. Zhu, P. Voigtlaender, H. Adam, B. Green, A. Geiger, B. Leibe, D. Cremers, A. Osep, L. Leal-Taixé, and L.-C. Chen, “STEP: Segmenting and tracking every pixel,” NeurIPS , 2021
2021
Earlier work this paper cites.
X. Li, L. Zhang, G. Cheng, K. Yang, Y. Tong, X. Zhu, and T. Xiang, “Global aggregation then local distribution for scene parsing,” IEEE TIP , 2021
2021
Earlier work this paper cites.
H. He, X. Li, G. Cheng, J. Shi, Y. Tong, G. Meng, V. Prinet, and L. Weng, “Enhanced boundary learning for glass-like object segmentation,” in ICCV , 2021
2021
Earlier work this paper cites.
S. Qiao, L.-C. Chen, and A. Yuille, “Detectors: Detecting objects with recursive feature pyramid and switchable atrous convolution,” in CVPR , 2021
2021
Earlier work this paper cites.
Y. Li, H. Zhao, X. Qi, L. Wang, Z. Li, J. Sun, and J. Jia, “Fully convolutional networks for panoptic segmentation,” CVPR , 2021
2021
Earlier work this paper cites.
Y. Fu, L. Yang, D. Liu, T. S. Huang, and H. Shi, “CompFeat: Comprehensive feature aggregation for video instance segmentation,” AAAI , 2021
2021
Earlier work this paper cites.
S. Qiao, Y. Zhu, H. Adam, A. Yuille, and L.-C. Chen, “Vip-deeplab: Learning visual perception with depth-aware video panoptic segmentation,” in CVPR , 2021
2021
Earlier work this paper cites.
R. Cheng, R. Razani, E. Taghavi, E. Li, and B. Liu, “(AF)2-S3Net: Attentive feature fusion with adaptive feature selection for sparse semantic segmentation network,” in CVPR , 2021
2021
Earlier work this paper cites.
Z. Zhou, Y. Zhang, and H. Foroosh, “Panoptic-PolarNet: Proposal-free lidar point cloud panoptic segmentation,” in CVPR , 2021
2021
Earlier work this paper cites.
F. Hong, H. Zhou, X. Zhu, H. Li, and Z. Liu, “LiDAR-based panoptic segmentation via dynamic shifting network,” in CVPR , 2021
2021
Earlier work this paper cites.
M. Aygun, A. Osep, M. Weber, M. Maximov, C. Stachniss, J. Behley, and L. Leal-Taixé, “4D panoptic lidar segmentation,” in CVPR , 2021
2021
Earlier work this paper cites.
X. Zhu, H. Zhou, T. Wang, F. Hong, Y. Ma, W. Li, H. Li, and D. Lin, “Cylindrical and asymmetrical 3d convolution networks for lidar segmentation,” in CVPR , 2021
2021
Earlier work this paper cites.
Z. Tian, C. Shen, H. Chen, and T. He, “FCOS: A simple and strong anchor-free object detector,” TPAMI , 2021
2021
Earlier work this paper cites.
E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo, “SegFormer: Simple and efficient design for semantic segmentation with transformers,” in NeurIPS , 2021
2021
Earlier work this paper cites.
S. Zheng, J. Lu, H. Zhao, X. Zhu, Z. Luo, Y. Wang, Y. Fu, J. Feng, T. Xiang, P. H. Torr, and L. Zhang, “Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers,” in CVPR , 2021
2021
Earlier work this paper cites.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in ICML , 2021
2021
Earlier work this paper cites.
H. Fan, B. Xiong, K. Mangalam, Y. Li, Z. Yan, J. Malik, and C. Feichtenhofer, “Multiscale vision transformers,” in ICCV , 2021
2021
Earlier work this paper cites.
A. Ali, H. Touvron, M. Caron, P. Bojanowski, M. Douze, A. Joulin, I. Laptev, N. Neverova, G. Synnaeve, J. Verbeek et al. , “XCiT: Cross-covariance image transformers,” NeurIPS , 2021
2021
Earlier work this paper cites.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” in ICCV , 2021
2021
Cited alongside, same era.
C.-F. R. Chen, Q. Fan, and R. Panda, “Crossvit: Cross-attention multi-scale vision transformer for image classification,” in ICCV , 2021
2021
Cited alongside, same era.
W. Xu, Y. Xu, T. Chang, and Z. Tu, “Co-scale conv-attentional image transformers,” in ICCV , 2021
2021
Cited alongside, same era.
X. Chu, Z. Tian, Y. Wang, B. Zhang, H. Ren, X. Wei, H. Xia, and C. Shen, “Twins: Revisiting the design of spatial attention in vision transformers,” in NeurIPS , 2021
2021
Cited alongside, same era.
H. Wu, B. Xiao, N. Codella, M. Liu, X. Dai, L. Yuan, and L. Zhang, “CvT: Introducing convolutions to vision transformers,” ICCV , 2021
2021
2022
Later among the works it cites.
K. Zhou, J. Yang, C. C. Loy, and Z. Liu, “Conditional prompt learning for vision-language models,” in CVPR , 2022
2022
Later among the works it cites.
R. Zhang, R. Fang, P. Gao, W. Zhang, K. Li, J. Dai, Y. Qiao, and H. Li, “Tip-Adapter: Training-free clip-adapter for better vision-language modeling,” ECCV , 2022
2022
Later among the works it cites.
Z. Lin, S. Geng, R. Zhang, P. Gao, G. de Melo, X. Wang, J. Dai, Y. Qiao, and H. Li, “Frozen clip models are efficient video learners,” in ECCV , 2022
2022
Later among the works it cites.
Y. Rao, W. Zhao, G. Chen, Y. Tang, Z. Zhu, G. Huang, J. Zhou, and J. Lu, “DenseCLIP: Language-guided dense prediction with context-aware prompting,” in CVPR , 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Y. Xu, Q. Zhang, J. Zhang, and D. Tao, “ViTAE: Vision transformer advanced by exploring intrinsic inductive bias,” NeurIPS , 2021
2021
Cited alongside, same era.
X. Chen*, S. Xie*, and K. He, “An empirical study of training self-supervised vision transformers,” ICCV , 2021
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in ICML , 2021
2021
Cited alongside, same era.
P. Sun, R. Zhang, Y. Jiang, T. Kong, C. Xu, W. Zhan, M. Tomizuka, L. Li, Z. Yuan, C. Wang, and P. Luo, “SparseR-CNN: End-to-end object detection with learnable proposals,” CVPR , 2021
2021
Cited alongside, same era.
Y. Fang, S. Yang, X. Wang, Y. Li, C. Fang, Y. Shan, B. Feng, and W. Liu, “Instances as queries,” in ICCV , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
B. Dong, F. Zeng, T. Wang, X. Zhang, and Y. Wei, “SOLQ: Segmenting objects by learning queries,” NeurIPS , 2021
2021
Cited alongside, same era.
2022
Later among the works it cites.
T. Lüddecke and A. Ecker, “Image segmentation using text and image prompts,” in CVPR , 2022
2022
Later among the works it cites.
X. Gu, T.-Y. Lin, W. Kuo, and Y. Cui, “Open-vocabulary object detection via vision and language knowledge distillation,” in ICLR , 2022
2022
Later among the works it cites.
X. Zhou, R. Girdhar, A. Joulin, P. Krähenbühl, and I. Misra, “Detecting twenty-thousand classes using image-level supervision,” in ECCV , 2022
2022
Later among the works it cites.
Y. Zang, W. Li, K. Zhou, C. Huang, and C. C. Loy, “Open-vocabulary detr with conditional matching,” in ECCV , 2022
2022
Later among the works it cites.
M. Maaz, H. Rasheed, S. Khan, F. S. Khan, R. M. Anwer, and M.-H. Yang, “Class-agnostic object detection with multi-modal transformer,” in ECCV , 2022
2022
Later among the works it cites.
G. Ghiasi, X. Gu, Y. Cui, and T.-Y. Lin, “Scaling open-vocabulary image segmentation with image-level labels,” in ECCV , 2022
2022
Later among the works it cites.
B. Li, K. Q. Weinberger, S. Belongie, V. Koltun, and R. Ranftl, “Language-driven semantic segmentation,” in ICLR , 2022
2022
Later among the works it cites.
A. Gupta, S. Narayan, K. Joseph, S. Khan, F. S. Khan, and M. Shah, “OW-DETR: Open-world detection transformer,” in CVPR , 2022
2022
Later among the works it cites.
L. Hoyer, D. Dai, and L. Van Gool, “DAFormer: Improving network architectures and training strategies for domain-adaptive semantic segmentation,” in CVPR , 2022
2022
Later among the works it cites.
——, “HRDA: Context-aware high-resolution domain-adaptive semantic segmentation,” in ECCV , 2022
2022
Later among the works it cites.
J. Yu, J. Liu, X. Wei, H. Zhou, Y. Nakata, D. Gudovskiy, T. Okuno, J. Li, K. Keutzer, and S. Zhang, “MTTrans: Cross-domain object detection with mean teacher transformer,” in ECCV , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
X. Zhou, V. Koltun, and P. Krähenbühl, “Simple multi-dataset detection,” in CVPR , 2022
2022
Later among the works it cites.
L. Xu, W. Ouyang, M. Bennamoun, F. Boussaid, and D. Xu, “Multi-class token transformer for weakly supervised semantic segmentation,” in CVPR , 2022
2022
Later among the works it cites.
S. Rossetti, D. Zappia, M. Sanzari, M. Schaerf, and F. Pirri, “Max pooling with vision transformers reconciles class and shape in weakly supervised semantic segmentation,” in ECCV , 2022
2022
Later among the works it cites.
J. Xu, S. De Mello, S. Liu, W. Byeon, T. Breuel, J. Kautz, and X. Wang, “Groupvit: Semantic segmentation emerges from text supervision,” in CVPR , 2022
2022
Later among the works it cites.
M. Hamilton, Z. Zhang, B. Hariharan, N. Snavely, and W. T. Freeman, “Unsupervised semantic segmentation by distilling feature correspondences,” ICLR , 2022
2022
Later among the works it cites.
G. Shin, W. Xie, and S. Albanie, “ReCo: Retrieve and co-segment for zero-shot transfer,” in NeurIPS , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
X. Wang, Z. Yu, S. De Mello, J. Kautz, A. Anandkumar, C. Shen, and J. M. Alvarez, “FreeSOLO: Learning to segment objects without annotations,” in CVPR , 2022
2022
Later among the works it cites.
M. Maaz, A. Shaker, H. Cholakkal, S. Khan, S. W. Zamir, R. M. Anwer, and F. Shahbaz Khan, “EdgeNeXt: efficiently amalgamated cnn-transformer architecture for mobile vision applications,” in ECCV Workshops , 2022
2022
Later among the works it cites.
S. Mehta and M. Rastegari, “Mobilevit: light-weight, general-purpose, and mobile-friendly vision transformer,” in ICLR , 2022
2022
Later among the works it cites.
W. Liang, Y. Yuan, H. Ding, X. Luo, W. Lin, D. Jia, Z. Zhang, C. Zhang, and H. Hu, “Expediting large-scale vision transformer for dense prediction without fine-tuning,” in NeurIPS , 2022
2022
Later among the works it cites.
W. Zhang, Z. Huang, G. Luo, T. Chen, X. Wang, W. Liu, G. Yu, and C. Shen, “TopFormer: Token pyramid transformer for mobile semantic segmentation,” in CVPR , 2022
2022
Later among the works it cites.
L. Ke, M. Danelljan, X. Li, Y.-W. Tai, C.-K. Tang, and F. Yu, “Mask transfiner for high-quality instance segmentation,” in CVPR , 2022
2022
Later among the works it cites.
L. Ke, H. Ding, M. Danelljan, Y.-W. Tai, C.-K. Tang, and F. Yu, “Video mask transfiner for high-quality video instance segmentation,” in ECCV , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
L. Qi, J. Kuen, Y. Wang, J. Gu, H. Zhao, P. Torr, Z. Lin, and J. Jia, “Open world entity segmentation,” TPAMI , 2022
2022
Later among the works it cites.
H. K. Cheng and A. G. Schwing, “XMem: Long-term video object segmentation with an atkinson-shiffrin memory model,” in ECCV , 2022
2022
Later among the works it cites.
K. Park, S. Woo, S. W. Oh, I. S. Kweon, and J.-Y. Lee, “Per-clip video object segmentation,” in CVPR , 2022
2022
Later among the works it cites.
H. Cao, Y. Wang, J. Chen, D. Jiang, X. Zhang, Q. Tian, and M. Wang, “Swin-Unet: Unet-like pure transformer for medical image segmentation,” in ECCV Workshops , 2022
2022
Later among the works it cites.
A. Hatamizadeh, Y. Tang, V. Nath, D. Yang, A. Myronenko, B. Landman, H. R. Roth, and D. Xu, “UNETR: Transformers for 3d medical image segmentation,” in WACV , 2022
2022
Later among the works it cites.
G. Sun, Y. Liu, H. Ding, T. Probst, and L. Van Gool, “Coarse-to-fine feature mining for video semantic segmentation,” in CVPR , 2022
2022
Later among the works it cites.
G. Sun, Y. Liu, H. Tang, A. Chhatkuli, L. Zhang, and L. Van Gool, “Mining relations among cross-frame affinities for video semantic segmentation,” ECCV , 2022
2022
Later among the works it cites.
Y. Zhou, H. Zhang, H. Lee, S. Sun, P. Li, Y. Zhu, B. Yoo, X. Qi, and J.-J. Han, “Slot-VPS: Object-centric representation learning for video panoptic segmentation,” in CVPR , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in CVPR , 2022
2022
Later among the works it cites.
J. Yang, Y. Z. Ang, Z. Guo, K. Zhou, W. Zhang, and Z. Liu, “Panoptic scene graph generation,” in ECCV , 2022
2022
Later among the works it cites.
X. Wu, Y. Lao, L. Jiang, X. Liu, and H. Zhao, “Point transformer v2: Grouped vector attention and partition-based pooling,” in NeurIPS , 2022
2022
Later among the works it cites.
J. Ding, N. Xue, G.-S. Xia, and D. Dai, “Decoupling zero-shot semantic segmentation,” in CVPR , 2022
2022
Later among the works it cites.
M. Xu, Z. Zhang, F. Wei, Y. Lin, Y. Cao, H. Hu, and X. Bai, “A simple baseline for open-vocabulary semantic segmentation with pre-trained vision-language model,” in ECCV , 2022
2022
Later among the works it cites.
C. Zhou, C. C. Loy, and B. Dai, “DenseCLIP: Extract free dense labels from clip,” ECCV , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
D. Huynh, J. Kuen, Z. Lin, J. Gu, and E. Elhamifar, “Open-vocabulary instance segmentation via robust cross-modal pseudo-labeling,” in CVPR , 2022
2022
Later among the works it cites.
M. Xu, Z. Zhang, F. Wei, Y. Lin, Y. Cao, H. Hu, and X. Bai, “A simple baseline for open-vocabulary semantic segmentation with pre-trained vision-language model,” in ECCV , 2022
2022
Later among the works it cites.
L. Ru, B. Du, Y. Zhan, and C. Wu, “Weakly-supervised semantic segmentation with visual words learning and hybrid pooling,” IJCV , 2022
2022
Later among the works it cites.
Y. Zhao, Z. Zhong, N. Zhao, N. Sebe, and G. H. Lee, “Style-hallucinated dual consistency learning for domain generalized semantic segmentation,” in ECCV , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
T. Zhou, F. Porikli, D. J. Crandall, L. Van Gool, and W. Wang, “A survey on deep learning technique for video segmentation,” TPAMI , 2023
2023
Closest in time.
C. Liu, H. Ding, and X. Jiang, “GRES: Generalized referring expression segmentation,” in CVPR , 2023
2023
Closest in time.
H. Ding, C. Liu, S. He, X. Jiang, P. H. Torr, and S. Bai, “MOSE: A new dataset for video object segmentation in complex scenes,” ICCV , 2023
2023
Closest in time.
H. Ding, C. Liu, S. He, X. Jiang, and C. C. Loy, “MeViS: A large-scale benchmark for video segmentation with motion expressions,” in ICCV , 2023
2023
Closest in time.
F. Li, A. Zeng, S. Liu, H. Zhang, H. Li, L. Zhang, and L. M. Ni, “Lite DETR: An interleaved multi-scale encoder for efficient detr,” CVPR , 2023
2023
Closest in time.
A. Athar, A. Hermans, J. Luiten, D. Ramanan, and B. Leibe, “TarViS: A unified approach for target-based video segmentation,” CVPR , 2023
2023
Closest in time.
2023
Closest in time.
H. Zhang, F. Li, S. Liu, L. Zhang, H. Su, J. Zhu, L. Ni, and H.-Y. Shum, “DINO: DETR with improved denoising anchor boxes for end-to-end object detection,” in ICLR , 2023
2023
Closest in time.
F. Li, H. Zhang, H. xu, S. Liu, L. Zhang, L. M. Ni, and H.-Y. Shum, “Mask DINO: Towards a unified transformer-based framework for object detection and segmentation,” in CVPR , 2023
2023
Closest in time.
D. Jia, Y. Yuan, H. He, X. Wu, H. Yu, W. Lin, L. Sun, C. Zhang, and H. Hu, “DETRs with hybrid matching,” in CVPR , 2023
2023
Closest in time.
X. Chen, M. Ding, X. Wang, Y. Xin, S. Mo, Y. Wang, S. Han, P. Luo, G. Zeng, and J. Wang, “Context autoencoder for self-supervised representation learning,” IJCV , 2023
2023
Closest in time.
K. Tian, Y. Jiang, Q. Diao, C. Lin, L. Wang, and Z. Yuan, “Designing bert for convolutional networks: Sparse and hierarchical masked modeling,” ICLR , 2023
2023
Closest in time.
X. Zou, Z.-Y. Dou, J. Yang, Z. Gan, L. Li, C. Li, X. Dai, J. Wang, L. Yuan, N. Peng, L. Wang, Y. J. Lee, and J. Gao, “Generalized decoding for pixel, image and language,” in CVPR , 2023
2023
Closest in time.
2023
Closest in time.
X. Li, H. Yuan, W. Zhang, J. Pang, G. Cheng, and C. C. Loy, “Tube-link: A flexible cross tube baseline for universal video segmentation,” ICCV , 2023
2023
Closest in time.
H. Zhang, F. Li, H. Xu, S. Huang, S. Liu, L. M. Ni, and L. Zhang, “MP-Former: Mask-piloted transformer for image segmentation,” CVPR , 2023
2023
Closest in time.
2023
Closest in time.
S. He, H. Ding, and W. Jiang, “Semantic-promoted debiasing and background disambiguation for zero-shot instance segmentation,” in CVPR , 2023
2023
Closest in time.
2023
Closest in time.
J. Schult, F. Engelmann, A. Hermans, O. Litany, S. Tang, and B. Leibe, “Mask3D for 3D Semantic Instance Segmentation,” in ICRA , 2023
2023
Closest in time.
J. Sun, C. Qing, J. Tan, and X. Xu, “Superpoint transformer for 3d scene instance segmentation,” AAAI , 2023
2023
Closest in time.
S. Su, J. Xu, H. Wang, Z. Miao, X. Zhan, D. Hao, and X. Li, “PUPS: Point cloud unified panoptic segmentation,” AAAI , 2023
2023
Closest in time.
Z. Chen, Y. Duan, W. Wang, J. He, T. Lu, J. Dai, and Y. Qiao, “Vision transformer adapter for dense predictions,” in ICLR , 2023
2023
Closest in time.
J. Jain, J. Li, M. Chiu, A. Hassani, N. Orlov, and H. Shi, “OneFormer: One transformer to rule universal image segmentation,” in CVPR , 2023
2023
Closest in time.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. Girshick, “Segment anything,” in ICCV , 2023
2023
Closest in time.
S. He, H. Ding, and W. Jiang, “Primitive generation and semantic-related alignment for universal zero-shot segmentation,” in CVPR , 2023
2023
Closest in time.
2023
Closest in time.
W. Kuo, Y. Cui, X. Gu, A. Piergiovanni, and A. Angelova, “F-VLM: Open-vocabulary object detection upon frozen vision and language models,” in ICLR , 2023
2023
Closest in time.
J. Wu, X. Li, H. Ding, X. Li, G. Cheng, Y. Tong, and C. C. Loy, “Betrayed by captions: Joint caption grounding and generation for open vocabulary instance segmentation,” in ICCV , 2023
2023
Closest in time.
J. Qin, J. Wu, P. Yan, M. Li, R. Yuxi, X. Xiao, Y. Wang, R. Wang, S. Wen, X. Pan et al. , “FreeSeg: Unified, universal and open-vocabulary image segmentation,” CVPR , 2023
2023
Closest in time.
J. Xu, S. Liu, A. Vahdat, W. Byeon, X. Wang, and S. De Mello, “Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion Models,” CVPR , 2023
2023
Closest in time.
M. Xu, Z. Zhang, F. Wei, H. Hu, and X. Bai, “Side adapter network for open-vocabulary semantic segmentation,” CVPR , 2023
2023
Closest in time.
L. Hoyer, D. Dai, H. Wang, and L. Van Gool, “MIC: Masked image consistency for context-enhanced domain adaptation,” in CVPR , 2023
2023
Closest in time.
B. Xie, S. Li, M. Li, C. H. Liu, G. Huang, and G. Wang, “Sepico: Semantic-guided pixel contrast for domain adaptive semantic segmentation,” PAMI , 2023
2023
Closest in time.
Z. Qiang, L. Yuang, L. Yuang, Y. Chaohui, L. Jingliang, W. Zhibin, and W. Fan, “LMSeg: Language-guided multi-dataset segmentation,” ICLR , 2023
2023
Closest in time.
L. Meng, X. Dai, Y. Chen, P. Zhang, D. Chen, M. Liu, J. Wang, Z. Wu, L. Yuan, and Y.-G. Jiang, “Detection hub: Unifying object detection datasets via query adaptation on language embedding,” CVPR , 2023
2023
Closest in time.
2023
Closest in time.
Y. Zhao, Z. Zhong, N. Zhao, N. Sebe, and G. H. Lee, “Style-hallucinated dual consistency learning: A unified framework for visual domain generalization,” IJCV , 2023
2023
Closest in time.
M. Yi, Q. Cui, H. Wu, C. Yang, O. Yoshie, and H. Lu, “A simple framework for text-supervised semantic segmentation,” in CVPR , 2023
2023
Closest in time.
L. Ke, M. Danelljan, H. Ding, Y.-W. Tai, C.-K. Tang, and F. Yu, “Mask-free video instance segmentation,” in CVPR , 2023
2023
Closest in time.
X. Wang, R. Girdhar, S. X. Yu, and I. Misra, “Cut and learn for unsupervised object detection and instance segmentation,” in CVPR , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Q. Wan, Z. Huang, J. Lu, G. Yu, and L. Zhang, “SeaFormer: Squeeze-enhanced axial transformer for mobile semantic segmentation,” in ICLR , 2023
2023
Closest in time.
M. Wang, H. Ding, J. H. Liew, J. Liu, Y. Zhao, and Y. Wei, “Segrefiner: Towards model-agnostic segmentation refinement with discrete diffusion process,” in NeurIPS , 2023
2023
Closest in time.
Q. Wen, J. Yang, X. Yang, and K. Liang, “Patchdct: Patch refinement for high quality instance segmentation,” ICLR , 2023
2023
Closest in time.
J. Wang, D. Chen, Z. Wu, C. Luo, C. Tang, X. Dai, Y. Zhao, Y. Xie, L. Yuan, and Y.-G. Jiang, “Look before you match: Instance understanding matters in video object segmentation,” in CVPR , 2023
2023
Closest in time.
H. Zhang, F. Li, X. Zou, S. Liu, C. Li, J. Yang, and L. Zhang, “A simple framework for open-vocabulary segmentation and detection,” in ICCV , 2023
2023
Closest in time.
2023
Closest in time.
X. Wang, S. Li, K. Kallidromitis, Y. Kato, K. Kozuka, and T. Darrell, “Hierarchical open-vocabulary universal image segmentation,” NeurIPS , 2023
2023
Closest in time.
J. Liang, T. Zhou, D. Liu, and W. Wang, “Clustseg: Clustering for universal segmentation,” ICML , 2023
2023
Closest in time.
J. Hu, L. Huang, T. Ren, S. Zhang, R. Ji, and L. Cao, “You only segment once: Towards real-time panoptic segmentation,” in CVPR , 2023
2023
Closest in time.
K. Ying, Q. Zhong, W. Mao, Z. Wang, H. Chen, L. Y. Wu, Y. Liu, C. Fan, Y. Zhuge, and C. Shen, “Ctvis: Consistent training for online video instance segmentation,” in ICCV , 2023
2023
Closest in time.
M. Heo, S. Hwang, J. Hyun, H. Kim, S. W. Oh, J.-Y. Lee, and S. J. Kim, “A generalized framework for video instance segmentation,” in CVPR , 2023
2023
Closest in time.
J. Lu, C. Clark, R. Zellers, R. Mottaghi, and A. Kembhavi, “Unified-IO: A unified model for vision, language, and multi-modal tasks,” ICLR , 2023
2023
Closest in time.
2023
Closest in time.
J. Yang, W. Peng, X. Li, Z. Guo, L. Chen, B. Li, Z. Ma, K. Zhou, W. Zhang, C. C. Loy, and Z. Liu, “Panoptic video scene graph generation,” in CVPR , 2023
2023
Closest in time.
J. Yang, J. CEN, W. PENG, S. Liu, F. Hong, X. Li, K. Zhou, Q. Chen, and Z. Liu, “4d panoptic scene graph generation,” in NeurIPS , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
J. Schult, F. Engelmann, A. Hermans, O. Litany, S. Tang, and B. Leibe, “Mask3d: Mask transformer for 3d semantic instance segmentation,” in ICRA , 2023
2023
Closest in time.
J. Sun, C. Qing, J. Tan, and X. Xu, “Superpoint transformer for 3d scene instance segmentation,” in AAAI , 2023
2023
Closest in time.
J. Wu, X. Li, S. Xu, H. Yuan, H. Ding, Y. Yang, X. Li, J. Zhang, Y. Tong, X. Jiang, B. Ghanem, and D. Tao, “Towards open vocabulary learning: A survey,” arXiv pre-print , 2023
2023
Closest in time.
Z. Ding, J. Wang, and Z. Tu, “Open-vocabulary panoptic segmentation with maskclip,” ICML , 2023
2023
Closest in time.
K. Han, Y. Liu, J. H. Liew, H. Ding, Y. Wei, J. Liu, Y. Wang, Y. Tang, Y. Yang, J. Feng et al. , “Global knowledge calibration for fast open-vocabulary segmentation,” ICCV , 2023
2023
Closest in time.
M. Xu, Z. Zhang, F. Wei, H. Hu, and X. Bai, “Side adapter network for open-vocabulary semantic segmentation,” CVPR , 2023
2023
Closest in time.
V. VS, N. Yu, C. Xing, C. Qin, M. Gao, J. C. Niebles, V. M. Patel, and R. Xu, “Mask-free ovis: Open-vocabulary instance segmentation without manual mask annotations,” CVPR , 2023
2023
Closest in time.
2023
Closest in time.
L. Ru, H. Zheng, Y. Zhan, and B. Du, “Token contrast for weakly-supervised semantic segmentation,” in CVPR , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
D. Bo, W. Pichao, and F. Wang, “Afformer: Head-free lightweight semantic segmentation with linear transformer,” in AAAI , 2023
2023
Closest in time.
2023
Closest in time.
X. Li, H. Yuan, W. Li, H. Ding, S. Wu, W. Zhang, Y. Li, K. Chen, and C. C. Loy, “Omg-seg: Is one model good enough for all segmentation?” in CVPR , 2024
2024
Closest in time.
S. He and H. Ding, “Decoupling static and hierarchical motion perception for referring video segmentation,” in CVPR , 2024
2024
Closest in time.
C. Liu, X. Li, and H. Ding, “Referring image editing: Object-level image editing via referring expressions,” in CVPR , 2024
2024
Closest in time.
S. Xu, H. Yuan, Q. Shi, L. Qi, J. Wang, Y. Yang, Y. Li, K. Chen, Y. Tong, B. Ghanem, X. Li, and M.-H. Yang, “Rap-sam: Towards real-time all-purpose segment anything,” arXiv preprint , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.