Fetching the paper…
Reading the bibliography…
Contrastive language-image pre-training (CLIP) is a powerful vision-language model that has shown great benefits for various tasks.
P. Chen, Q. Li, S. Biaz, T. Bui, A. Nguyen, gscorecam: What objects is clip looking at?, in: Proceedings of the Asian Conference on Computer Vision, 2022, pp. 1959–1975
1975
Earlier work this paper cites.
T.-S. Chua, J. Tang, R. Hong, H. Li, Z. Luo, Y. Zheng, Nus-wide: a real-world web image database from national university of singapore, in: Proceedings of the ACM international conference on image and video retrieval, 2009, pp. 1–9
2009
Earlier work this paper cites.
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, A. Zisserman, The pascal visual object classes (voc) challenge, International journal of computer vision 88 (2) (2010) 303–338
2010
Earlier work this paper cites.
R. Mottaghi, X. Chen, X. Liu, N.-G. Cho, S.-W. Lee, S. Fidler, R. Urtasun, A. Yuille, The role of context for object detection and semantic segmentation in the wild, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2014, pp. 891–898
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, C. L. Zitnick, Microsoft coco: Common objects in context, in: European conference on computer vision, Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, A. Torralba, Learning deep features for discriminative localization, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 2921–2929
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778
2016
Earlier work this paper cites.
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, B. Schiele, The cityscapes dataset for semantic urban scene understanding, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 3213–3223
2016
Earlier work this paper cites.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, D. Batra, Grad-cam: Visual explanations from deep networks via gradient-based localization, in: Proceedings of the IEEE international conference on computer vision, 2017, pp. 618–626
2017
Earlier work this paper cites.
H. Caesar, J. Uijlings, V. Ferrari, Coco-stuff: Thing and stuff classes in context, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 1209–1218
2018
Earlier work this paper cites.
P. Sharma, N. Ding, S. Goodman, R. Soricut, Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning, in: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2018, pp. 2556–2565
2018
Earlier work this paper cites.
S. Lapuschkin, S. Wäldchen, A. Binder, G. Montavon, W. Samek, K.-R. Müller, Unmasking clever hans predictors and assessing what machines really learn, Nature communications 10 (1) (2019) 1096
2019
Earlier work this paper cites.
Y. Li, Z. Kuang, L. Liu, Y. Chen, W. Zhang, Pseudo-mask matters in weakly-supervised semantic segmentation, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 6964–6973
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al., Learning transferable visual models from natural language supervision, in: International Conference on Machine Learning, PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
H. Chefer, S. Gur, L. Wolf, Transformer interpretability beyond attention visualization, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 782–791
2021
Earlier work this paper cites.
H. Chefer, S. Gur, L. Wolf, Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 397–406
2021
Cited alongside, same era.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al., An image is worth 16x16 words: Transformers for image recognition at scale (2021)
2021
Cited alongside, same era.
O. Bar-Tal, D. Ofri-Amar, R. Fridman, Y. Kasten, T. Dekel, Text2live: Text-driven layered image and video editing, in: Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XV, Springer, 2022, pp. 707–723
2022
Cited alongside, same era.
M. X. et al., A simple baseline for open vocabulary semantic segmentation with pre-trained vision-language model, Proceedings of the IEEE/CVF European Conference on Computer Vision (ECCV) (2022)
2022
Z. Tan, X. Yang, Z. Ye, Q. Wang, Y. Yan, A. Nguyen, K. Huang, Semantic similarity distance: Towards better text-image consistency metric in text-to-image generation, Pattern Recognition 144 (2023) 109883
2023
Closest in time.
Y. Huang, A. Jia, X. Zhang, J. Zhang, Generic attention-model explainability by weighted relevance accumulation, in: Proceedings of the 5th ACM International Conference on Multimedia in Asia, 2023, pp. 1–7
2023
Closest in time.
H. Zhang, F. Li, X. Zou, S. Liu, C. Li, J. Yang, L. Zhang, A simple framework for open-vocabulary segmentation and detection, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 1020–1031
2023
Closest in time.
J. Xu, J. Hou, Y. Zhang, R. Feng, Y. Wang, Y. Qiao, W. Xie, Learning open-vocabulary semantic segmentation models from natural language supervision, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 2935–2944
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2022
Cited alongside, same era.
R. Paiss, H. Chefer, L. Wolf, No token left behind: Explainability-aided image classification and generation, in: European Conference on Computer Vision, Springer, 2022, pp. 334–350
2022
Cited alongside, same era.
J. Xu, S. De Mello, S. Liu, W. Byeon, T. Breuel, J. Kautz, X. Wang, Groupvit: Semantic segmentation emerges from text supervision, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 18134–18144
2022
Cited alongside, same era.
C. Zhou, C. C. Loy, B. Dai, Extract free dense labels from clip, in: Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXVIII, Springer, 2022, pp. 696–712
2022
Cited alongside, same era.
2022
Cited alongside, same era.
X. Sun, P. Hu, K. Saenko, Dualcoop: Fast adaptation to multi-label recognition with limited annotations, Advances in Neural Information Processing Systems 35 (2022) 30569–30582
2022
Cited alongside, same era.
S. Gao, Z.-Y. Li, M.-H. Yang, M.-M. Cheng, J. Han, P. Torr, Large-scale unsupervised semantic segmentation, IEEE Transactions on Pattern Analysis and Machine Intelligence (2022)
2022
Cited alongside, same era.
G. Shin, W. Xie, S. Albanie, Reco: Retrieve and co-segment for zero-shot transfer, Advances in Neural Information Processing Systems 35 (2022) 33754–33767
2022
Cited alongside, same era.
Closest in time.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al., Segment anything, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 4015–4026
2023
Closest in time.
Z. Guo, B. Dong, Z. Ji, J. Bai, Y. Guo, W. Zuo, Texts as images in prompt tuning for multi-label image recognition, in: CVPR, 2023, pp. 2808–2817
2023
Closest in time.
H. Luo, J. Bao, Y. Wu, X. He, T. Li, Segclip: Patch aggregation with learnable centers for open-vocabulary semantic segmentation, in: International Conference on Machine Learning, PMLR, 2023, pp. 23033–23044
2023
Closest in time.
J. Chen, D. Zhu, G. Qian, B. Ghanem, Z. Yan, C. Zhu, F. Xiao, S. C. Culatana, M. Elhoseiny, Exploring open-vocabulary semantic segmentation from clip vision encoder distillation only, in: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023, pp. 699–710
2023
Closest in time.
S. He, T. Guo, T. Dai, R. Qiao, X. Shu, B. Ren, S.-T. Xia, Open-vocabulary multi-label classification via multi-modal knowledge transfer, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37, 2023, pp. 808–816
2023
Closest in time.
M. Ô. V. Ngc, E. Carlinet, J. Fabrizio, T. Géraud, The dahu graph-cut for interactive segmentation on 2d/3d images, Pattern Recognition 136 (2023) 109207
2023
Closest in time.
L. Baur, K. Ditschuneit, M. Schambach, C. Kaymakci, T. Wollmann, A. Sauer, Explainability and interpretability in electric load forecasting using machine learning techniques–a review, Energy and AI (2024) 100358
2024
Closest in time.
H. Yu, H. Lu, M. Zhao, Z. Li, G. Gu, Gradient aggregation based fine-grained image retrieval: A unified viewpoint for cnn and transformer, Pattern Recognition 149 (2024) 110248
2024
Closest in time.
F. Zheng, J. Cao, W. Yu, Z. Chen, N. Xiao, Y. Lu, Exploring low-resource medical image classification with weakly supervised prompt learning, Pattern Recognition 149 (2024) 110250
2024
Closest in time.
Y. Li, H. Liang, H. Zheng, R. Yu, Cr-cam: Generating explanations for deep neural networks by contrasting and ranking features, Pattern Recognition 149 (2024) 110251
2024
Closest in time.
R. Huang, X. Pan, H. Zheng, H. Jiang, Z. Xie, C. Wu, S. Song, G. Huang, Joint representation learning for text and 3d point cloud, Pattern Recognition 147 (2024) 110086
2024
Closest in time.