Fetching the paper…
Reading the bibliography…
Recently, the emergence of the large-scale vision-language model (VLM), such as CLIP, has opened the way towards open-world object perception.
Everingham, M., Van Gool, L., Williams, C.K., Winn, J., Zisserman, A.: The pascal visual object classes (voc) challenge. International Journal of Computer Vision (2010)
2010
Earlier work this paper cites.
Katsuki, F., Constantinidis, C.: Bottom-up and top-down attention: different processes and overlapping neural systems. The Neuroscientist (2014)
2014
Earlier work this paper cites.
Mottaghi, R., Chen, X., Liu, X., Cho, N.G., Lee, S.W., Fidler, S., Urtasun, R., Yuille, A.: The role of context for object detection and semantic segmentation in the wild. In: Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (2014)
2014
Earlier work this paper cites.
Bideau, P., Learned-Miller, E.: It’s moving! a probabilistic model for causal motion segmentation in moving camera videos. In: Proceedings of European Conference on Computer Vision (2016)
2016
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (2016)
2016
Earlier work this paper cites.
Lin, D., Dai, J., Jia, J., He, K., Sun, J.: Scribblesup: Scribble-supervised convolutional networks for semantic segmentation. Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (2016)
2016
Earlier work this paper cites.
Litjens, G., Kooi, T., Bejnordi, B.E., Setio, A.A.A., Ciompi, F., Ghafoorian, M., van der Laak, J.A., van Ginneken, B., Sánchez, C.I.: A survey on deep learning in medical image analysis. Medical Image Analysis (2017)
2017
Earlier work this paper cites.
Neuhold, G., Ollmann, T., Bulo, S.R., Kontschieder, P.: The mapillary vistas dataset for semantic understanding of street scenes. In: Proceedings of the IEEE International Conference on Computer Vision (2017)
2017
Earlier work this paper cites.
Skurowski, P., Abdulameer, H., Błaszczyk, J., Depta, T., Kornacki, A., Kozieł, P.: Animal camouflage analysis: Chameleon database (2017), http://kgwisc.aei.polsl.pl/index.php/pl/dataset/63-animal-camouflage-analysis
2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, u., Polosukhin, I.: Attention is all you need. In: International Conference on Neural Information Processing Systems (2017)
2017
Earlier work this paper cites.
Zhao, H., Puig, X., Zhou, B., Fidler, S., Torralba, A.: Open vocabulary scene parsing. In: Proceedings of the IEEE International Conference on Computer Vision (2017)
2017
Earlier work this paper cites.
Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., Torralba, A.: Scene parsing through ade20k dataset. In: Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (2017)
2017
Earlier work this paper cites.
Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., Zhang, L.: Bottom-up and top-down attention for image captioning and visual question answering. In: Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (2018)
2018
Earlier work this paper cites.
Caesar, H., Uijlings, J., Ferrari, V.: Coco-stuff: Thing and stuff classes in context. In: Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (2018)
2018
Earlier work this paper cites.
Mithun, N.C., Panda, R., Papalexakis, E.E., Roy-Chowdhury, A.K.: Webly supervised joint embedding for cross-modal image-text retrieval. In: Proceedings of the ACM International Conference on Multimedia (2018)
2018
Earlier work this paper cites.
Le, T.N., Nguyen, T.V., Nie, Z., Tran, M.T., Sugimoto, A.: Anabranch network for camouflaged object segmentation. Computer Vision and Image Understanding (2019)
2019
Earlier work this paper cites.
Liu, L., Wang, R., Xie, C., Yang, P., Wang, F., Sudirman, S., Liu, W.: Pestnet: An end-to-end deep learning approach for large-scale multi-class pest detection and classification. IEEE Access (2019)
2019
Earlier work this paper cites.
Loshchilov, I., Hutter, F.: Decoupled weight decay regularization. In: International Conference on Learning Representations (2019)
2019
Earlier work this paper cites.
McCrae, J.P., Rademaker, A., Fellbaum, C.D.: English wordnet 2019 – an open-source wordnet for english. In: Global WordNet Conference (2019)
2019
Earlier work this paper cites.
Su, W., Zhu, X., Cao, Y., Li, B., Lu, L., Wei, F., Dai, J.: Vl-bert: Pre-training of generic visual-linguistic representations. In: International Conference on Learning Representations (2019)
2019
Earlier work this paper cites.
Zheng, Y., Zhang, X., Wang, F., Cao, T., Sun, M., Wang, X.: Detection of people with camouflage pattern via dense deconvolution network. IEEE Signal Processing Letters (Jan 2019)
2019
Earlier work this paper cites.
Caesar, H., Bankiti, V., Lang, A.H., Vora, S., Liong, V.E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., Beijbom, O.: nuscenes: A multimodal dataset for autonomous driving. In: Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (2020)
2020
Earlier work this paper cites.
Chen, Y.C., Li, L., Yu, L., El Kholy, A., Ahmed, F., Gan, Z., Cheng, Y., Liu, J.: Uniter: Universal image-text representation learning. In: Proceedings of European Conference on Computer Vision (2020)
2020
Earlier work this paper cites.
Fan, D.P., Ji, G.P., Sun, G., Cheng, M.M., Shen, J., Shao, L.: Camouflaged object detection. In: Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (2020)
2020
Earlier work this paper cites.
Fan, D.P., Zhou, T., Ji, G.P., Zhou, Y., Chen, G., Fu, H., Shen, J., Shao, L.: Inf-net: Automatic covid-19 lung infection segmentation from ct images. IEEE Transactions on Medical Imaging (2020)
2020
Earlier work this paper cites.
He, K., Gkioxari, G., Dollar, P., Girshick, R.: Mask r-cnn. IEEE Transactions on Pattern Analysis and Machine Intelligence (Feb 2020)
2020
Earlier work this paper cites.
Ji, W., Li, J., Zhang, M., Piao, Y., Lu, H.: Accurate rgb-d salient object detection via collaborative learning. In: Proceedings of European Conference on Computer Vision (2020)
2020
Earlier work this paper cites.
Li, G., Duan, N., Fang, Y., Gong, M., Jiang, D.: Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training. In: AAAI Conference on Artificial Intelligence (2020)
2020
Earlier work this paper cites.
Li, X., Yin, X., Li, C., Zhang, P., Hu, X., Zhang, L., Wang, L., Hu, H., Dong, L., Wei, F., et al.: Oscar: Object-semantics aligned pre-training for vision-language tasks. In: Proceedings of European Conference on Computer Vision (2020)
2020
Earlier work this paper cites.
Pang, Y., Zhang, L., Zhao, X., Lu, H.: Hierarchical dynamic filtering network for rgb-d salient object detection. In: Proceedings of European Conference on Computer Vision (2020)
2020
Earlier work this paper cites.
Pang, Y., Zhao, X., Zhang, L., Lu, H.: Multi-scale interactive network for salient object detection. In: Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (2020)
2020
Cited alongside, same era.
Zhang, M., Ji, W., Piao, Y., Li, J., Zhang, Y., Xu, S., Lu, H.: Lfnet: Light field fusion network for salient object detection. IEEE Transactions on Image Processing (2020)
2020
Cited alongside, same era.
Zhao, X., Pang, Y., Zhang, L., Lu, H., Zhang, L.: Suppress and balance: A simple gated network for salient object detection. In: Proceedings of European Conference on Computer Vision (2020)
2020
Cited alongside, same era.
Zhao, X., Zhang, L., Pang, Y., Lu, H., Zhang, L.: A single stream network for robust and real-time rgb-d salient object detection. In: Proceedings of European Conference on Computer Vision (2020)
2020
Cited alongside, same era.
Unal, O., Dai, D., Gool, L.V.: Scribble-supervised lidar semantic segmentation. Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (2022)
2022
Later among the works it cites.
Xiang, M., Zhang, J., Lv, Y., Li, A., Zhong, Y., Dai, Y.: Exploring depth contribution for camouflaged object detection (2022)
2022
Later among the works it cites.
Yin, B., Zhang, X., Hou, Q., Sun, B.Y., Fan, D.P., Van Gool, L.: Camoformer: Masked separable attention for camouflaged object detection (2022)
2022
Later among the works it cites.
Zhao, X., Pang, Y., Zhang, L., Lu, H.: Joint learning of salient object detection, depth estimation and contour extraction. IEEE Transactions on Image Processing (2022)
2022
Later among the works it cites.
Zhao, X., Pang, Y., Zhang, L., Lu, H., Ruan, X.: Self-supervised pretraining for rgb-d salient object detection. In: AAAI Conference on Artificial Intelligence (2022)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cheng, B., Schwing, A., Kirillov, A.: Per-pixel classification is not all you need for semantic segmentation. International Conference on Neural Information Processing Systems (2021)
2021
Cited alongside, same era.
Fan, D.P., Ji, G.P., Cheng, M.M., Shao, L.: Concealed object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence (2021)
2021
Cited alongside, same era.
Gu, X., Lin, T.Y., Kuo, W., Cui, Y.: Open-vocabulary object detection via vision and language knowledge distillation. In: International Conference on Learning Representations (2021)
2021
Cited alongside, same era.
Li, A., Zhang, J., Lyu, Y., Liu, B., Zhang, T., Dai, Y.: Uncertainty-aware joint salient object and camouflaged object detection. In: Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (2021)
2021
Cited alongside, same era.
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., Guo, B.: Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE International Conference on Computer Vision (2021)
2021
Cited alongside, same era.
Lyu, Y., Zhang, J., Dai, Y., Li, A., Liu, B., Barnes, N., Fan, D.P.: Simultaneously localize, segment and rank the camouflaged objects. In: Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (2021)
2021
Cited alongside, same era.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: Proceedings of the International Conference on Machine Learning (2021)
2021
Cited alongside, same era.
Ranftl, R., Bochkovskiy, A., Koltun, V.: Vision transformers for dense prediction. In: Proceedings of the IEEE International Conference on Computer Vision (2021)
2021
Cited alongside, same era.
2022
Later among the works it cites.
Zhou, C., Loy, C.C., Dai, B.: Extract free dense labels from clip. In: Proceedings of European Conference on Computer Vision (2022)
2022
Later among the works it cites.
Cho, S., Shin, H., Hong, S., An, S., Lee, S., Arnab, A., Seo, P.H., Kim, S.W.: Cat-seg: Cost aggregation for open-vocabulary semantic segmentation. arXiv preprint (2023)
2023
Closest in time.
Fan, D.P., Ji, G.P., Xu, P., Cheng, M.M., Sakaridis, C., Van Gool, L.: Advances in deep concealed scene understanding. Visual Intelligence (2023)
2023
Closest in time.
Ji, W., Li, J., Bian, C., Zhou, Z., Zhao, J., Yuille, A.L., Cheng, L.: Multispectral video semantic segmentation: A benchmark dataset and baseline. In: Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (2023)
2023
Closest in time.
Li, J., Ji, W., Wang, S., Li, W., Cheng, L.: Dvsod: Rgb-d video salient object detection. In: International Conference on Neural Information Processing Systems (2023)
2023
Closest in time.
Li, J., Ji, W., Zhang, M., Piao, Y., Lu, H., Cheng, L.: Delving into calibrated depth for accurate rgb-d salient object detection. International Journal of Computer Vision (2023)
2023
Closest in time.
Pang, Y., Zhao, X., Zhang, L., Lu, H.: Caver: Cross-modal view-mixed transformer for bi-modal salient object detection. IEEE Transactions on Image Processing (2023)
2023
Closest in time.
Rizzo, M., Marcuzzo, M., Zangari, A., Gasparetto, A., Albarelli, A.: Fruit ripeness classification: A survey (2023)
2023
Closest in time.
Thisanke, H., Deshan, C., Chamith, K., Seneviratne, S., Vidanaarachchi, R., Herath, D.: Semantic segmentation using vision transformers: A survey. arXiv preprint (2023)
2023
Closest in time.
Xu, J., Liu, S., Vahdat, A., Byeon, W., Wang, X., De Mello, S.: Open-vocabulary panoptic segmentation with text-to-image diffusion models. In: Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (2023)
2023
Closest in time.
Xu, M., Zhang, Z., Wei, F., Hu, H., Bai, X.: Side adapter network for open-vocabulary semantic segmentation. Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (2023)
2023
Closest in time.
Yang, J.: Plantcamo dataset (2023), https://github.com/yjybuaa/PlantCamo
2023
Closest in time.
Yang, L., Xu, X., Kang, B., Shi, Y., Zhao, H.: Freemask: Synthetic images with dense annotations make stronger segmentation models. In: International Conference on Neural Information Processing Systems (2023)
2023
Closest in time.
Ye, H., Kuen, J., Liu, Q., Lin, Z., Price, B., Xu, D.: Seggen: Supercharging segmentation models with text2mask and mask2img synthesis. ArXiv (2023)
2023
Closest in time.
Yu, Q., He, J., Deng, X., Shen, X., Chen, L.C.: Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip. In: International Conference on Neural Information Processing Systems (2023)
2023
Closest in time.
Zabari, N., Hoshen, Y.: Open-vocabulary semantic segmentation using test-time distillation. In: European Conference on Computer Vision Workshops (2023)
2023
Closest in time.
Zhang, M., Yao, S., Hu, B., Piao, Y., Ji, W.: C2dfnet: Criss-cross dynamic filter network for rgb-d salient object detection. IEEE Transactions on Multimedia (2023)
2023
Closest in time.
2023
Closest in time.
Zhu, C., Chen, L.: A survey on open-vocabulary detection and segmentation: Past, present, and future. arXiv preprint (2023)
2023
Closest in time.
Pang, Y., Zhao, X., Xiang, T.Z., Zhang, L., Lu, H.: Zoomnext: A unified collaborative pyramid network for camouflaged object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)
2024
Closest in time.
Wu, J., Li, X., Xu, S., Yuan, H., Ding, H., Yang, Y., Li, X., Zhang, J., Tong, Y., Jiang, X., Ghanem, B., Tao, D.: Towards open vocabulary learning: A survey. arXiv preprint (2024)
2024
Closest in time.
Yu, Q., Zhao, X., Pang, Y., Zhang, L., Lu, H.: Multi-view aggregation network for dichotomous image segmentation. In: Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (2024)
2024
Closest in time.
Zhao, X., Chang, S., Pang, Y., Yang, J., Zhang, L., Lu, H.: Multi-source fusion and automatic predictor selection for zero-shot video object segmentation. International Journal of Computer Vision (2024)
2024
Closest in time.
Zhao, X., Pang, Y., Ji, W., Sheng, B., Zuo, J., Zhang, L., Lu, H.: Spider: A unified framework for context-dependent concept understanding. In: Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (2024)
2024
Closest in time.
Zhao, X., Pang, Y., Zhang, L., Lu, H., Zhang, L.: Towards diverse binary segmentation via a simple yet general gated network. International Journal of Computer Vision (2024)
2024
Closest in time.