Fetching the paper…
Reading the bibliography…
In the field of visual scene understanding, deep neural networks have made impressive advancements in various core tasks like segmentation, tracking, and detection.
H. W. Kuhn, “The hungarian method for the assignment problem,” Naval research logistics quarterly , 1955
1955
Earlier work this paper cites.
P. F. Felzenszwalb and D. P. Huttenlocher, “Efficient graph-based image segmentation,” IJCV , 2004
2004
Earlier work this paper cites.
V. Ordonez, G. Kulkarni, and T. Berg, “Im2text: Describing images using 1 million captioned photographs,” NeurIPS , 2011
2011
Earlier work this paper cites.
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean, “Distributed representations of words and phrases and their compositionality,” NeurIPS , 2013
2013
Earlier work this paper cites.
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean, “Distributed representations of words and phrases and their compositionality,” in NeurIPS , 2013
2013
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in ECCV , 2014
2014
Earlier work this paper cites.
A. Bendale and T. E. Boult, “Towards open world recognition,” CVPR , 2014
2014
Earlier work this paper cites.
R. Mottaghi, X. Chen, X. Liu, N.-G. Cho, S.-W. Lee, S. Fidler, R. Urtasun, and A. Yuille, “The role of context for object detection and semantic segmentation in the wild,” in CVPR , 2014
2014
Earlier work this paper cites.
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg, “Referitgame: Referring to objects in photographs of natural scenes,” in EMNLP , 2014
2014
Earlier work this paper cites.
B. Romera-Paredes and P. Torr, “An embarrassingly simple approach to zero-shot learning,” in ICML , 2015
2015
Earlier work this paper cites.
A. Bendale and T. Boult, “Towards open world recognition,” in CVPR , 2015
2015
Earlier work this paper cites.
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick, “Microsoft coco captions: Data collection and evaluation server,” CVPR , 2015
2015
Earlier work this paper cites.
M. Everingham, S. A. Eslami, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes challenge: A retrospective,” IJCV , 2015
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” NeurIPS , 2015
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “Imagenet large scale visual recognition challenge,” IJCV , 2015
2015
Earlier work this paper cites.
Z. Wu, S. Song, A. Khosla, F. Yu, L. Zhang, X. Tang, and J. Xiao, “3d shapenets: A deep representation for volumetric shapes,” in CVPR , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016
2016
Earlier work this paper cites.
Y. Xian, Z. Akata, G. Sharma, Q. Nguyen, M. Hein, and B. Schiele, “Latent embeddings for zero-shot classification,” in CVPR , 2016
2016
Earlier work this paper cites.
A. Bendale and T. E. Boult, “Towards open set deep networks,” CVPR , 2016
2016
Earlier work this paper cites.
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” in CVPR , 2016
2016
Earlier work this paper cites.
D. Hendrycks and K. Gimpel, “A baseline for detecting misclassified and out-of-distribution examples in neural networks,” ICLR , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
G. Patterson and J. Hays, “Coco attributes: Attributes for people, animals, and objects,” in ECCV , 2016
2016
Earlier work this paper cites.
R. R. Selvaraju, A. Das, R. Vedantam, M. Cogswell, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” IJCV , 2016
2016
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in ICCV , 2017
2017
Earlier work this paper cites.
L. Shu, H. Xu, and B. Liu, “Doc: Deep open classification of text documents,” ACL , 2017
2017
Earlier work this paper cites.
H. Zhao, X. Puig, B. Zhou, S. Fidler, and A. Torralba, “Open vocabulary scene parsing,” in ICCV , 2017
2017
Earlier work this paper cites.
B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba, “Semantic understanding of scenes through the ADE20K dataset,” CVPR , 2017
2017
Earlier work this paper cites.
N. Kardan and K. O. Stanley, “Mitigating fooling with competitive overcomplete output layer neural networks,” IJCNN , 2017
2017
Earlier work this paper cites.
Z. Ge, S. Demyanov, Z. Chen, and R. Garnavi, “Generative openmax for multi-class open set classification,” BMVC , 2017
2017
Earlier work this paper cites.
Y. Yu, W.-Y. Qu, N. Li, and Z. Guo, “Open category classification by adversarial sample generation,” in IJCAI , 2017
2017
Earlier work this paper cites.
K. Lee, H. Lee, K. Lee, and J. Shin, “Training confidence-calibrated classifiers for detecting out-of-distribution samples,” ICLR , 2017
2017
Earlier work this paper cites.
C. Peng, X. Zhang, G. Yu, G. Luo, and J. Sun, “Large kernel matters–improve semantic segmentation by global convolutional network,” in CVPR , 2017
2017
Earlier work this paper cites.
H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid scene parsing network,” in CVPR , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in ICCV , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
R. Yeh, J. Xiong, W.-M. Hwu, M. Do, and A. Schwing, “Interpretable and globally optimal prediction for textual grounding using image concepts,” NeurIPS , 2017
2017
Earlier work this paper cites.
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in CVPR , 2017
2017
Earlier work this paper cites.
N. Wojke, A. Bewley, and D. Paulus, “Simple online and realtime tracking with a deep association metric,” in ICIP , 2017
2017
Earlier work this paper cites.
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner, “Scannet: Richly-annotated 3d reconstructions of indoor scenes,” in CVPR , 2017
2017
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma et al. , “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” IJCV , 2017
2017
Earlier work this paper cites.
A. Bansal, K. Sikka, G. Sharma, R. Chellappa, and A. Divakaran, “Zero-shot object detection,” in ECCV , 2018
2018
Earlier work this paper cites.
T. Baltrušaitis, C. Ahuja, and L.-P. Morency, “Multimodal machine learning: A survey and taxonomy,” TPAMI , 2018
2018
Earlier work this paper cites.
A. Bansal, K. Sikka, G. Sharma, R. Chellappa, and A. Divakaran, “Zero-shot object detection,” in ECCV , 2018
2018
Earlier work this paper cites.
C. Geng, S.-J. Huang, and S. Chen, “Recent advances in open set recognition: A survey,” TPAMI , 2018
2018
Earlier work this paper cites.
A. R. Dhamija, M. Günther, and T. E. Boult, “Reducing network agnostophobia,” NeurIPS , 2018
2018
Earlier work this paper cites.
M. Hassen and P. K. Chan, “Learning a neural-network-based representation for open set recognition,” in SDM , 2018
2018
Earlier work this paper cites.
L. Neal, M. L. Olson, X. Z. Fern, W.-K. Wong, and F. Li, “Open set learning with counterfactual images,” in ECCV , 2018
2018
Earlier work this paper cites.
I. Jo, J. Kim, H. Kang, Y.-D. Kim, and S. Choi, “Open set recognition by regularising classifier with fake data generated by generative adversarial networks,” ICASSP , 2018
2018
Earlier work this paper cites.
C. Geng and S. Chen, “Collective decision for open set recognition,” TKDE , 2018
2018
Earlier work this paper cites.
K. Lee, K. Lee, K. Min, Y. Zhang, J. Shin, and H. Lee, “Hierarchical novelty detection for visual object recognition,” CVPR , 2018
2018
Earlier work this paper cites.
B. Zong, Q. Song, M. R. Min, W. Cheng, C. Lumezanu, D. ki Cho, and H. Chen, “Deep autoencoding gaussian mixture model for unsupervised anomaly detection,” in ICLR , 2018
2018
Earlier work this paper cites.
D. Abati, A. Porrello, S. Calderara, and R. Cucchiara, “Latent space autoregression for novelty detection,” CVPR , 2018
2018
Earlier work this paper cites.
S. Pidhorskyi, R. Almohsen, D. A. Adjeroh, and G. Doretto, “Generative probabilistic novelty detection with adversarial autoencoders,” in NeurIPS , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
K. Rakelly, E. Shelhamer, T. Darrell, A. Efros, and S. Levine, “Conditional networks for few-shot semantic segmentation,” ICLR , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
C. Yu, J. Wang, C. Peng, C. Gao, G. Yu, and N. Sang, “Learning a discriminative feature network for semantic segmentation,” in CVPR , 2018
2018
Earlier work this paper cites.
H. Ding, X. Jiang, B. Shuai, A. Q. Liu, and G. Wang, “Context contrasted feature and gated multi-scale aggregation for scene segmentation,” in CVPR , 2018
2018
Earlier work this paper cites.
X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in CVPR , 2018
2018
Earlier work this paper cites.
H. Hu, J. Gu, Z. Zhang, J. Dai, and Y. Wei, “Relation networks for object detection,” in CVPR , 2018
2018
Earlier work this paper cites.
B. Graham, M. Engelcke, and L. Van Der Maaten, “3D semantic segmentation with submanifold sparse convolutional networks,” in CVPR , 2018
2018
Earlier work this paper cites.
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in ACL , 2018
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in ACL , 2019
2019
Earlier work this paper cites.
W. Wang, V. W. Zheng, H. Yu, and C. Miao, “A survey of zero-shot learning: Settings, methods, and applications,” ACM TIST , 2019
2019
Earlier work this paper cites.
A. Gupta, P. Dollar, and R. Girshick, “Lvis: A dataset for large vocabulary instance segmentation,” in CVPR , 2019
2019
Earlier work this paper cites.
L. Yang, Y. Fan, and N. Xu, “Video instance segmentation,” in ICCV , 2019
2019
Earlier work this paper cites.
Y. Yang, C. Hou, Y. Lang, D. Guan, D. Huang, and J. Xu, “Open-set human activity recognition based on micro-doppler signatures,” PR , 2019
2019
Earlier work this paper cites.
Y. Xian, S. Choudhury, Y. He, B. Schiele, and Z. Akata, “Semantic projection network for zero-and few-label semantic segmentation,” in CVPR , 2019
2019
Earlier work this paper cites.
M. Bucher, T.-H. Vu, M. Cord, and P. Pérez, “Zero-shot semantic segmentation,” NeurIPS , 2019
2019
Earlier work this paper cites.
X. Yan, Z. Chen, A. Xu, X. Wang, X. Liang, and L. Lin, “Meta r-cnn: Towards general solver for instance-level low-shot learning,” in ICCV , 2019
2019
Earlier work this paper cites.
H. Ding, X. Jiang, A. Q. Liu, N. M. Thalmann, and G. Wang, “Boundary-aware feature propagation for scene segmentation,” in ICCV , 2019
2019
Earlier work this paper cites.
J. Fu, J. Liu, H. Tian, Z. Fang, and H. Lu, “Dual attention network for scene segmentation,” in CVPR , 2019
2019
Earlier work this paper cites.
Z. Tian, C. Shen, H. Chen, and T. He, “Fcos: Fully convolutional one-stage object detection,” in ICCV , 2019
2019
Earlier work this paper cites.
D. Neven, B. D. Brabandere, M. Proesmans, and L. V. Gool, “Instance segmentation by jointly optimizing spatial embeddings and clustering bandwidth,” in CVPR , 2019
2019
Earlier work this paper cites.
K. Chen, J. Pang, J. Wang, Y. Xiong, X. Li, S. Sun, W. Feng, Z. Liu, J. Shi, W. Ouyang, C. C. Loy, and D. Lin, “Hybrid task cascade for instance segmentation,” in CVPR , 2019
2019
Earlier work this paper cites.
D. Bolya, C. Zhou, F. Xiao, and Y. J. Lee, “Yolact: Real-time instance segmentation,” in ICCV , 2019
2019
Earlier work this paper cites.
Y. Wu, A. Kirillov, F. Massa, W.-Y. Lo, and R. Girshick, “Detectron2,” https://github.com/facebookresearch/detectron2 , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” NeurIPS , 2019
2019
Earlier work this paper cites.
H. Tan and M. Bansal, “Lxmert: Learning cross-modality encoder representations from transformers,” ACL , 2019
2019
Earlier work this paper cites.
M. A. Uy, Q.-H. Pham, B.-S. Hua, T. Nguyen, and S.-K. Yeung, “Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 1588–1597
2019
Earlier work this paper cites.
S. Shao, Z. Li, T. Zhang, C. Peng, G. Yu, X. Zhang, J. Li, and J. Sun, “Objects365: A large-scale, high-quality dataset for object detection,” in ICCV , 2019
2019
Earlier work this paper cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in ECCV , 2020
2020
Earlier work this paper cites.
Y. Wang, Q. Yao, J. T. Kwok, and L. M. Ni, “Generalizing from a few examples: A survey on few-shot learning,” ACM Computing Surveys , 2020
2020
Earlier work this paper cites.
A. Dave, T. Khurana, P. Tokmakov, C. Schmid, and D. Ramanan, “Tao: A large-scale benchmark for tracking any object,” in ECCV , 2020
2020
Earlier work this paper cites.
Y.-C. Hsu, Y. Shen, H. Jin, and Z. Kira, “Generalized odin: Detecting out-of-distribution image without learning from out-of-distribution data,” CVPR , 2020
2020
Earlier work this paper cites.
E. Techapanurak, M. Suganuma, and T. Okatani, “Hyperparameter-free out-of-distribution detection using cosine similarity,” in ACCV , 2020
2020
Earlier work this paper cites.
Z. Gu, S. Zhou, L. Niu, Z. Zhao, and L. Zhang, “Context-aware feature generation for zero-shot semantic segmentation,” in ACM MM , 2020
2020
Earlier work this paper cites.
P. Li, Y. Wei, and Y. Yang, “Consistent structural relation learning for zero-shot segmentation,” NeurIPS , 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
G. Tian, S. Wang, J. Feng, L. Zhou, and Y. Mu, “Cap2seg: Inferring semantic and spatial context from captions for zero-shot image segmentation,” in ACM MM , 2020
2020
Earlier work this paper cites.
J. Liu, Y. Sun, C. Han, Z. Dou, and W. Li, “Deep representation learning on long-tailed data: A learnable embedding augmentation perspective,” in CVPR , 2020
2020
Earlier work this paper cites.
J. Ren, C. Yu, X. Ma, H. Zhao, S. Yi et al. , “Balanced meta-softmax for long-tailed visual recognition,” NeurIPS , 2020
2020
Earlier work this paper cites.
Y. Li, T. Wang, B. Kang, S. Tang, C. Wang, J. Li, and J. Feng, “Overcoming classifier imbalance for long-tail object detection with balanced group softmax,” in CVPR , 2020
2020
Earlier work this paper cites.
X. Hu, Y. Jiang, K. Tang, J. Chen, C. Miao, and H. Zhang, “Learning to segment the tail,” in CVPR , 2020
2020
Earlier work this paper cites.
T. Wang, Y. Li, B. Kang, J. Li, J. Liew, S. Tang, S. Hoi, and J. Feng, “The devil is in classification: A simple framework for long-tail instance segmentation,” in ECCV , 2020
2020
Earlier work this paper cites.
X. Wang, T. Huang, J. Gonzalez, T. Darrell, and F. Yu, “Frustratingly simple few-shot object detection,” in ICML , 2020
2020
Earlier work this paper cites.
Z. Fan, J.-G. Yu, Z. Liang, J. Ou, C. Gao, G.-S. Xia, and Y. Li, “Fgn: Fully guided network for few-shot instance segmentation,” in CVPR , 2020
2020
Earlier work this paper cites.
H. Ding, X. Jiang, B. Shuai, A. Q. Liu, and G. Wang, “Semantic segmentation with context encoding and multi-path decoding,” TIP , 2020
2020
Earlier work this paper cites.
X. Li, H. Zhao, L. Han, Y. Tong, S. Tan, and K. Yang, “Gated fully fusion for semantic segmentation,” in AAAI , 2020
2020
Earlier work this paper cites.
X. Li, A. You, Z. Zhu, H. Zhao, M. Yang, K. Yang, and Y. Tong, “Semantic flow for fast and accurate scene parsing,” in ECCV , 2020
2020
Earlier work this paper cites.
A. Kirillov, Y. Wu, K. He, and R. Girshick, “Pointrend: Image segmentation as rendering,” in CVPR , 2020
2020
Earlier work this paper cites.
X. Li, X. Li, L. Zhang, G. Cheng, J. Shi, Z. Lin, S. Tan, and Y. Tong, “Improving semantic segmentation via decoupled body and edge supervision,” in ECCV , 2020
2020
Earlier work this paper cites.
Z. Tian, C. Shen, and H. Chen, “Conditional convolutions for instance segmentation,” in ECCV , 2020
2020
Earlier work this paper cites.
R. Zhang, Z. Tian, C. Shen, M. You, and Y. Yan, “Mask encoding for single shot instance segmentation,” in CVPR , 2020
2020
Earlier work this paper cites.
B. Cheng, M. D. Collins, Y. Zhu, T. Liu, T. S. Huang, H. Adam, and L.-C. Chen, “Panoptic-deeplab: A simple, strong, and fast baseline for bottom-up panoptic segmentation,” in CVPR , 2020
2020
Earlier work this paper cites.
X. Wang, R. Zhang, T. Kong, L. Li, and C. Shen, “SOLOv2: Dynamic and fast instance segmentation,” in NeurIPS , 2020
2020
Earlier work this paper cites.
H. Wang, Y. Zhu, B. Green, H. Adam, A. Yuille, and L.-C. Chen, “Axial-deeplab: Stand-alone axial-attention for panoptic segmentation,” in ECCV , 2020
2020
Earlier work this paper cites.
G. Li, N. Duan, Y. Fang, M. Gong, and D. Jiang, “Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training,” in AAAI , 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
G. Luo, Y. Zhou, X. Sun, L. Cao, C. Wu, C. Deng, and R. Ji, “Multi-task collaborative network for joint referring expression comprehension and segmentation,” in CVPR , 2020
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” NeurIPS , 2020
2020
Earlier work this paper cites.
S. Zhang, C. Chi, Y. Yao, Z. Lei, and S. Z. Li, “Bridging the gap between anchor-based and anchor-free detection via adaptive training sample selection,” CVPR , 2020
2020
Earlier work this paper cites.
H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom, “nuscenes: A multimodal dataset for autonomous driving,” in CVPR , 2020
2020
Earlier work this paper cites.
C. Wu, Z. Lin, S. Cohen, T. Bui, and S. Maji, “Phrasecut: Language-based image segmentation in the wild,” in CVPR , 2020
2020
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR , 2021
2021
Cited alongside, same era.
A. Zareian, K. D. Rosa, D. H. Hu, and S.-F. Chang, “Open-vocabulary object detection using captions,” CVPR , 2021
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in ICML , 2021
2021
Cited alongside, same era.
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” in ICML , 2021
2021
Cited alongside, same era.
H. Ding, C. Liu, S. He, X. Jiang, P. H. Torr, and S. Bai, “MOSE: A new dataset for video object segmentation in complex scenes,” in ICCV , 2023
2023
Closest in time.
2023
Closest in time.
S. He, H. Ding, and W. Jiang, “Semantic-promoted debiasing and background disambiguation for zero-shot instance segmentation,” in CVPR , 2023
2023
Closest in time.
S. He, X. Jiang, W. Jiang, and H. Ding, “Prototype adaption and projection for few- and zero-shot 3d point cloud semantic segmentation,” TIP , 2023
2023
Closest in time.
S. He, H. Ding, and W. Jiang, “Primitive generation and semantic-related alignment for universal zero-shot segmentation,” in CVPR , 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Köhler, M. Eisenbach, and H.-M. Gross, “Few-shot object detection: A comprehensive survey,” TNNLS , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” ICCV , 2021
2021
Cited alongside, same era.
K. Joseph, S. Khan, F. S. Khan, and V. N. Balasubramanian, “Towards open world object detection,” in CVPR , 2021
2021
Cited alongside, same era.
K. J. Joseph, S. H. Khan, F. S. Khan, and V. N. Balasubramanian, “Towards open world object detection,” CVPR , 2021
2021
Cited alongside, same era.
Y. Zheng, J. Wu, Y. Qin, F. Zhang, and L. Cui, “Zero-shot instance segmentation,” in CVPR , 2021
2021
Cited alongside, same era.
J. Miao, Y. Wei, Y. Wu, C. Liang, G. Li, and Y. Yang, “Vspw: A large-scale dataset for video scene parsing in the wild,” in CVPR , 2021
2021
Cited alongside, same era.
2023
Closest in time.
H. Ding, H. Zhang, and X. Jiang, “Self-regularized prototypical network for few-shot semantic segmentation,” PR , 2023
2023
Closest in time.
S. Wu, W. Zhang, S. Jin, W. Liu, and C. C. Loy, “Aligning bag of regions for open-vocabulary object detection,” in CVPR , 2023
2023
Closest in time.
X. Chen, S. Li, S.-N. Lim, A. Torralba, and H. Zhao, “Open-vocabulary panoptic segmentation with embedding modulation,” ICCV , 2023
2023
Closest in time.
S. Peng, K. Genova, C. M. Jiang, A. Tagliasacchi, M. Pollefeys, and T. Funkhouser, “Openscene: 3d scene understanding with open vocabularies,” in CVPR , 2023
2023
Closest in time.
L. Xue, M. Gao, C. Xing, R. Martín-Martín, J. Wu, C. Xiong, R. Xu, J. C. Niebles, and S. Savarese, “Ulip: Learning unified representation of language, image and point cloud for 3d understanding,” in CVPR , 2023
2023
Closest in time.
2023
Closest in time.
F. Liang, B. Wu, X. Dai, K. Li, Y. Zhao, H. Zhang, P. Zhang, P. Vajda, and D. Marculescu, “Open-vocabulary semantic segmentation with mask-adapted clip,” CVPR , 2023
2023
Closest in time.
W. Lin, L. Karlinsky, N. Shvetsova, H. Possegger, M. Kozinski, R. Panda, R. Feris, H. Kuehne, and H. Bischof, “Match, expand and improve: Unsupervised finetuning for zero-shot action recognition with language knowledge,” in ICCV , 2023
2023
Closest in time.
R. Ding, J. Yang, C. Xue, W. Zhang, S. Bai, and X. Qi, “PLA: Language-driven open-vocabulary 3d scene understanding,” in CVPR , 2023
2023
Closest in time.
H. Luo, J. Bao, Y. Wu, X. He, and T. Li, “Segclip: Patch aggregation with learnable centers for open-vocabulary semantic segmentation,” in ICML , 2023
2023
Closest in time.
V. VS, N. Yu, C. Xing, C. Qin, M. Gao, J. C. Niebles, V. M. Patel, and R. Xu, “Mask-free ovis: Open-vocabulary instance segmentation without manual mask annotations,” CVPR , 2023
2023
Closest in time.
H. Zhang, F. Li, X. Zou, S. Liu, C. Li, J. Gao, J. Yang, and L. Zhang, “A simple framework for open-vocabulary segmentation and detection,” ICCV , 2023
2023
Closest in time.
2023
Closest in time.
X. Zou, Z.-Y. Dou, J. Yang, Z. Gan, L. Li, C. Li, X. Dai, H. Behl, J. Wang, L. Yuan et al. , “Generalized decoding for pixel, image, and language,” CVPR , 2023
2023
Closest in time.
J. Qin, J. Wu, P. Yan, M. Li, R. Yuxi, X. Xiao, Y. Wang, R. Wang, S. Wen, X. Pan et al. , “Freeseg: Unified, universal and open-vocabulary image segmentation,” CVPR , 2023
2023
Closest in time.
2023
Closest in time.
J. Xu, S. Liu, A. Vahdat, W. Byeon, X. Wang, and S. De Mello, “Open-vocabulary panoptic segmentation with text-to-image diffusion models,” CVPR , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
W. Kuo, Y. Cui, X. Gu, A. J. Piergiovanni, and A. Angelova, “F-vlm: Open-vocabulary object detection upon frozen vision and language models,” ICLR , 2023
2023
Closest in time.
2023
Closest in time.
W. Wu, Y. Zhao, M. Z. Shou, H. Zhou, and C. Shen, “Diffumask: Synthesizing images with pixel-level annotations for semantic segmentation using diffusion models,” ICCV , 2023
2023
Closest in time.
C. Liu, H. Ding, and X. Jiang, “GRES: Generalized referring expression segmentation,” in CVPR , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Z. Hu, Y. Sun, J. Wang, and Y. Yang, “DAC-DETR: Divide the attention layers and conquer,” in NeurIPS , 2023
2023
Closest in time.
Y. Li, H. Fan, R. Hu, C. Feichtenhofer, and K. He, “Scaling language-image pre-training via masking,” in CVPR , 2023
2023
Closest in time.
2023
Closest in time.
J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,” in ICML , 2023
2023
Closest in time.
H. Ding, C. Liu, S. He, X. Jiang, and C. C. Loy, “MeViS: A large-scale benchmark for video segmentation with motion expressions,” in ICCV , 2023
2023
Closest in time.
M. Li, C. Wang, W. Feng, S. Lyu, G. Cheng, X. Li, B. Liu, and Q. Zhao, “Iterative robust visual grounding with masked reference based centerpoint supervision,” in ICCVW , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
X. Wu, F. Zhu, R. Zhao, and H. Li, “Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching,” CVPR , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
L. Wang, Y. Liu, P. Du, Z. Ding, Y. Liao, Q. Qi, B. Chen, and S. Liu, “Object-aware distillation pyramid for open-vocabulary object detection,” CVPR , 2023
2023
Closest in time.
2023
Closest in time.
K. Chen, X. Jiang, Y. Hu, X. Tang, Y. Gao, J. Chen, and W. Xie, “Ovarnet: Towards open-vocabulary object attribute recognition,” CVPR , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
P. Kaul, W. Xie, and A. Zisserman, “Multi-modal classifiers for open-vocabulary object detection,” ICML , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
L. Yao, J. Han, X. Liang, D. Xu, W. Zhang, Z. Li, and H. Xu, “Detclipv2: Scalable open-vocabulary object detection pre-training via word-region alignment,” CVPR , 2023
2023
Closest in time.
C. Ma, Y. Jiang, X. Wen, Z. Yuan, and X. Qi, “Codet: Co-occurrence guided region-word alignment for open-vocabulary object detection,” in NeurIPS , 2023
2023
Closest in time.
D. Kim, A. Angelova, and W. Kuo, “Region-aware pretraining for open-vocabulary object detection with vision transformers,” CVPR , 2023
2023
Closest in time.
M. Xu, Z. Zhang, F. Wei, H. Hu, and X. Bai, “Side adapter network for open-vocabulary semantic segmentation,” CVPR , 2023
2023
Closest in time.
2023
Closest in time.
K. Han, Y. Liu, J. H. Liew, H. Ding, Y. Wei, J. Liu, Y. Wang, Y. Tang, Y. Yang, J. Feng et al. , “Global knowledge calibration for fast open-vocabulary segmentation,” ICCV , 2023
2023
Closest in time.
X. Chen, S. Li, S.-N. Lim, A. Torralba, and H. Zhao, “Open-vocabulary panoptic segmentation with embedding modulation,” ICCV , 2023
2023
Closest in time.
2023
Closest in time.
Q. Yu, J. He, X. Deng, X. Shen, and L.-C. Chen, “Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip,” in NeurIPS , 2023
2023
Closest in time.
2023
Closest in time.
J. Mukhoti, T.-Y. Lin, O. Poursaeed, R. Wang, A. Shah, P. H. Torr, and S.-N. Lim, “Open vocabulary semantic segmentation with patch aligned contrastive learning,” in CVPR , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Z. Li, Q. Zhou, X. Zhang, Y. Zhang, Y. Wang, and W. Xie, “Guiding text-to-image diffusion model towards grounded generation,” ICCV , 2023
2023
Closest in time.
H. Rasheed, M. U. Khattak, M. Maaz, S. Khan, and F. S. Khan, “Fine-tuned clip models are efficient video learners,” in CVPR , 2023
2023
Closest in time.
S. Li, T. Fischer, L. Ke, H. Ding, M. Danelljan, and F. Yu, “Ovtrack: Open-vocabulary multiple object tracking,” in CVPR , 2023
2023
Closest in time.
2023
Closest in time.
X. Zhu, R. Zhang, B. He, Z. Zeng, S. Zhang, and P. Gao, “Pointclip v2: Adapting clip for powerful 3d open-world learning,” in ICCV , 2023
2023
Closest in time.
Y. Zeng, C. Jiang, J. Mao, J. Han, C. Ye, Q. Huang, D.-Y. Yeung, Z. Yang, X. Liang, and H. Xu, “Clip2: Contrastive language-image-point pretraining from real-world point cloud data,” in CVPR , 2023
2023
Closest in time.
2023
Closest in time.
Z. Weng, X. Yang, A. Li, Z. Wu, and Y.-G. Jiang, “Transforming clip to an open-vocabulary video model via interpolated weight optimization,” in ICML , 2023
2023
Closest in time.
T. Yang, Y. Zhu, Y. Xie, A. Zhang, C. Chen, and M. Li, “AIM: Adapting image models for efficient video action recognition,” in ICLR , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
T. Huang, B. Dong, Y. Yang, X. Huang, R. W. Lau, W. Ouyang, and W. Zuo, “Clip2point: Transfer clip to point cloud classification with image-depth pre-training,” in ICCV , 2023
2023
Closest in time.
Y. Lu, C. Xu, X. Wei, X. Xie, M. Tomizuka, K. Keutzer, and S. Zhang, “Open-vocabulary point-cloud object detection without 3d annotation,” in CVPR , 2023
2023
Closest in time.
Y. Cao, Y. Zeng, H. Xu, and D. Xu, “Coda: Collaborative novel box discovery and cross-modal alignment for open-vocabulary 3d object detection,” in NeurIPS , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
J. Zhang, R. Dong, and K. Ma, “Clip-fo3d: Learning free open-world 3d scene representations from 2d dense clip,” arXiv pre-print , 2023
2023
Closest in time.
2023
Closest in time.
M. Liu, Y. Zhu, H. Cai, S. Han, Z. Ling, F. Porikli, and H. Su, “PartSLIP: Low-shot part segmentation for 3d point clouds via pretrained image-language models,” in CVPR , 2023
2023
Closest in time.
2023
Closest in time.
A. Takmaz, E. Fedele, R. W. Sumner, M. Pollefeys, F. Tombari, and F. Engelmann, “Openmask3d: Open-vocabulary 3d instance segmentation,” in NeurIPS , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
R. Chen, Y. Liu, L. Kong, X. Zhu, Y. Ma, Y. Li, Y. Hou, Y. Qiao, and W. Wang, “Clip2scene: Towards label-efficient 3d scene understanding by clip,” in CVPR , 2023
2023
Closest in time.
L. Qi, J. Kuen, W. Guo, T. Shen, J. Gu, W. Li, J. Jia, Z. Lin, and M.-H. Yang, “Fine-grained entity segmentation,” ICCV , 2023
2023
Closest in time.
O. Zohar, K.-C. Wang, and S. Yeung, “Prob: Probabilistic objectness for open world object detection,” CVPR , 2023
2023
Closest in time.
P. Sun, S. Chen, C. Zhu, F. Xiao, P. Luo, S. Xie, and Z. Yan, “Going denser with open-vocabulary part segmentation,” ICCV , 2023
2023
Closest in time.
L. Qi, J. Kuen, W. Guo, J. Gu, Z. Lin, B. Du, Y. Xu, and M.-H. Yang, “Aims: All-inclusive multi-level segmentation,” NeurIPS , 2023
2023
Closest in time.
T.-Y. Pan, Q. Liu, W.-L. Chao, and B. Price, “Towards open-world segmentation of parts,” in CVPR , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Z. Fang, X. Li, X. Li, J. M. Buhmann, C. C. Loy, and M. Liu, “Explore in-context learning for 3d point cloud understanding,” NeurIPS , 2023
2023
Closest in time.
X. Wang, W. Wang, Y. Cao, C. Shen, and T. Huang, “Images speak in images: A generalist painter for in-context visual learning,” CVPR , 2023
2023
Closest in time.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” NeurIPS , 2023
2023
Closest in time.
W. Wu, Z. Sun, and W. Ouyang, “Revisiting classifier: Transferring vision-language models for video recognition,” in AAAI , 2023
2023
Closest in time.
M. Deitke, D. Schwenk, J. Salvador, L. Weihs, O. Michel, E. VanderBilt, L. Schmidt, K. Ehsani, A. Kembhavi, and A. Farhadi, “Objaverse: A universe of annotated 3d objects,” in CVPR , 2023
2023
Closest in time.
2023
Closest in time.
C. Shi and S. Yang, “Edadet: Open-vocabulary object detection using early dense alignment,” ICCV , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
X. Wang, S. Li, K. Kallidromitis, Y. Kato, K. Kozuka, and T. Darrell, “Hierarchical open-vocabulary universal image segmentation,” NeurIPS , 2023
2023
Closest in time.
2023
Closest in time.
C. Pham, T. Vu, and K. Nguyen, “Lp-ovod: Open-vocabulary object detection by linear probing,” WACV , 2024
2024
Closest in time.
X. Li, H. Yuan, W. Li, H. Ding, S. Wu, W. Zhang, Y. Li, K. Chen, and C. C. Loy, “Omg-seg: Is one model good enough for all segmentation?” arXiv , 2024
2024
Closest in time.
G. Hess, A. Tonderski, C. Petersson, K. Åström, and L. Svensson, “Lidarclip or: How i learned to talk to point clouds,” in WACV , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. Jeong, G. Park, J. Yoo, H. Jung, and H. Kim, “Proxydet: Synthesizing proxy novel classes via classwise mixup for open-vocabulary object detection,” AAAI , 2024
2024
Closest in time.